Model watch · Trend watch
From the Aug 2, 2026 daily brief
Overnight's 15 new papers are still unprocessed and this column has no single-day increment, so here is a claims-versus-academia pairing. Vinyals's self-reported changed mind from the past year: training on narrow domains where answers auto-verify, like math and coding (what the AI world calls RLVR, reinforcement learning with verifiable rewards), "creates this generalization. I think that is not something I quite predicted to work as well as it … did"; but he says plainly the question is open — is math-and-code alone enough to induce general problem-solving? "I don't know. I mean, I think it could go either way." His bet on the way through is letting the model be the judge, to open up domains that can't be auto-scored (Unsupervised Learning interview, recorded May 22, 2026). The academic measurements land squarely on that bet: a July 2026 evaluation study finds AI judges to be systematically over-lenient when no reference answer is available, with verdicts flipping in up to 85% of cases once reference answers are supplied (arXiv paper, July 2026); a June 2026 grader-design study finds that unit tests used as graders can have zero detection power for specific error types — and that a wrong grader's feedback is worse than no feedback at all (arXiv paper, June 2026). Side by side: the domains where you can't write a grader are exactly the domains where AI judging is currently least reliable. The bet isn't unplayable — it's that the measurements aren't on its side yet.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
A newer route in AI training is rubric-based rewards: instead of training a separate rewar…
Epoch AI, the AI data-research organization, and the evaluator METR released MirrorCode, a…
The current mainstay of post-training (the stage after base training where a model is taug…
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…