SecondSourceJudgment rebuilt from primary sources
Research · Jul 18, 2026

The gap between "scoring high on competition math" and "actually being able to derive" now has hard numbers.

Research notes · Evidence update (original paper March 2026)

From the Jul 18, 2026 daily brief

A Meta FAIR-affiliated team's March paper, Principia, argues that current math benchmarks mostly grade a final value or a multiple-choice pick, while real scientific work demands deriving formulas and structure. They built a 2,558-problem derivation benchmark. At the time we only recorded the qualitative claim that frontier models struggle; this week we read the PDF directly and filled in the numbers (arXiv, Mar 2026): OpenAI's o3 scores 62.90 on the derivation benchmark, while the same model scores 85.63 on competition math AIME-2024 — a gap of about 23 points. The most memorable result is a separate experiment, run on the math-and-engineering subset of the existing SuperGPQA benchmark, not on the 2,558 problems themselves: take the same questions, merely remove the answer options, require derivation instead — and strong models drop roughly 6 to 14 points. o3 falls from 69.10 to 62.90; Qwen3-235B from 69.33 to 55.58. Some fraction of those high scores came from answer-choice cues rather than derivation. For anyone using benchmark scores as a selling point or acceptance criterion, the actionable takeaway: never take a single headline number — demand the control condition, the same questions with the options stripped, or what you are buying may be a score propped up by answer-choice cues. Caveat: the grader is o3 itself — the self-built-eval, self-reported-results bias still applies.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section