SecondSourceJudgment rebuilt from primary sources
Research · Jul 19, 2026

Letting the model teach itself: majority voting upgraded from answer filter to verbatim textbook.

Research notes · This week (paper Jul 15)

From the Jul 19, 2026 daily brief

For problems with no answer key, sample many solutions from the model, take the majority-consensus answer — then train the model verbatim on the solutions that reach that answer. The team (Gkountouras, Jukić, Titov) self-reports: on the authors' chosen math-reasoning benchmarks (pass@1), gains of up to 12 points; with about one-seventh the compute it beats label-free reinforcement learning (training the model to self-adjust by trial and error) by 6 points; and after training the model solves problems it had failed in 32 straight prior attempts. That last claim is aimed squarely at the old objection that self-training only makes the model more confident about what it already knows (arXiv, Jul 15). A single self-report, no third-party reproduction; the "AI improving itself without human labels" research line stays on our watch list.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section