Models · This month Trend (published July 2026)
From the Jul 26, 2026 daily brief
The current mainstay of post-training (the stage after base training where a model is taught how to answer) is RLVR — reinforcement learning with verifiable rewards, which rewards the model only on tasks with checkable answers, like math and code. Its structural blind spot: it optimizes only what can be objectively scored and is numb to style, structure, and the hard-to-articulate qualities that matter — hence the failures everyone has seen: output turns monotone and unnatural, and models learn to game the score rather than write well. An MIT team's July paper proposes a fix: add an adversarial discriminator that learns to tell "human-written" from "model-generated," feeding that hard-to-quantify quality signal back into training. Self-reported results: on bug-fixing, markedly smaller edits at unchanged scores; on story-writing, output more human-like and more diverse (arXiv, Jul 2026). A single team's self-report, not independently replicated — directional for now. It is the other face of today's thread 6: the dimensions verifiable rewards cannot reach are still bridged by human demonstration — which is exactly the mechanistic reason the expert-data business exists. No other major new papers within this week's scan.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…
The Jacobian conjecture, posed in 1939, is a famous problem in algebraic geometry. It says…
Our July 21 Research Notes covered this empirical study (gains from optimizing an agent pi…
Alex Zhang, author of the RLM framework paper, argues that carefully designed task orchest…