Model watch · Trend watch (posts originally published July 17)
From the Aug 3, 2026 daily brief
First, which yardstick produced "one point apart": the composite intelligence index from the third-party evaluator Artificial Analysis — a leaderboard that weights multiple tests into a single total. On that board, Kimi K3 and Anthropic's previous-generation flagship Opus 4.8 sit one point apart. The actual readings: K3 at 57, Opus 4.8 at 56, with GPT-5.6 at 59 and Fable 5 at 60 on the same board — K3 ranks third, and the "one point" happens inside a band where the leaders bunch within four points of each other. As for the board's full-scale calibration and per-test weighting, they are not in our library; all we can report is the same-board scores and ranks — we cannot decompose how the total is built. Teortaxes then spreads out the per-task results: K3 matches Opus 4.8 on nearly every task, and the only significant lag is a single math-integrals item; raise that item to DeepSeek's 90.7, and the total would reach 78.8 (Teortaxes, July 17). A caution here: the 90.7 and 78.8 come from a different third-party per-task evaluation, which Teortaxes's post attributes to a third party called Lisan. What Lisan is — an individual evaluator, a tool, or an institution — this brief could not confirm, and the original leaderboard was never obtained; those two numbers trace back only to Teortaxes's relay. That evaluation and the intelligence index above are two different yardsticks whose scales don't interoperate — so those numbers can't be subtracted against the index's one-point gap, and putting them side by side doesn't mean they measure the same thing. Our handling: the per-task board is not in our library, the summation is Teortaxes's own arithmetic, this brief has not recomputed it, and the two numbers stand unverified. His judgment in passing is worth more than the table: "math-maxxing seems to have ceased being a priority" — pushing math to its extreme is exiting the frontier labs' resource allocation (same thread). Verification: Single-source table reading, original board not in hand; "math is exiting" runs against the reality that labs still showcase competition math loudly — a possible reconciliation is that showcasing and training priority are two different things; we keep the tension and don't reconcile it. Judgment update: "K3 can substitute for Opus 4.8" needs a reframe: the two capability profiles genuinely overlap at the per-task level — this is not a case of offsetting strengths that happen to cancel — so stop comparing on a single composite index when selecting models; look at where the per-task residuals fall. For product builders, one operational line: a product that needs strong math cannot expect the next general-purpose flagship to improve it automatically — pull the math item out and measure it separately when you select.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
Two lines from the K3 launch materials (verbatim, as relayed by Teortaxes): late in develo…
Overnight's 15 new papers are still unprocessed and this column has no single-day incremen…
A newer route in AI training is rubric-based rewards: instead of training a separate rewar…
Epoch AI, the AI data-research organization, and the evaluator METR released MirrorCode, a…