SecondSourceJudgment rebuilt from primary sources
Research · Aug 3, 2026

The "one point apart" between K3 and Opus 4.8, taken down to task level: the capability profiles genuinely overlap, and the residual concentrates in one math item.

Model watch · Trend watch (posts originally published July 17)

From the Aug 3, 2026 daily brief

First, which yardstick produced "one point apart": the composite intelligence index from the third-party evaluator Artificial Analysis — a leaderboard that weights multiple tests into a single total. On that board, Kimi K3 and Anthropic's previous-generation flagship Opus 4.8 sit one point apart. The actual readings: K3 at 57, Opus 4.8 at 56, with GPT-5.6 at 59 and Fable 5 at 60 on the same board — K3 ranks third, and the "one point" happens inside a band where the leaders bunch within four points of each other. As for the board's full-scale calibration and per-test weighting, they are not in our library; all we can report is the same-board scores and ranks — we cannot decompose how the total is built. Teortaxes then spreads out the per-task results: K3 matches Opus 4.8 on nearly every task, and the only significant lag is a single math-integrals item; raise that item to DeepSeek's 90.7, and the total would reach 78.8 (Teortaxes, July 17). A caution here: the 90.7 and 78.8 come from a different third-party per-task evaluation, which Teortaxes's post attributes to a third party called Lisan. What Lisan is — an individual evaluator, a tool, or an institution — this brief could not confirm, and the original leaderboard was never obtained; those two numbers trace back only to Teortaxes's relay. That evaluation and the intelligence index above are two different yardsticks whose scales don't interoperate — so those numbers can't be subtracted against the index's one-point gap, and putting them side by side doesn't mean they measure the same thing. Our handling: the per-task board is not in our library, the summation is Teortaxes's own arithmetic, this brief has not recomputed it, and the two numbers stand unverified. His judgment in passing is worth more than the table: "math-maxxing seems to have ceased being a priority" — pushing math to its extreme is exiting the frontier labs' resource allocation (same thread). Verification: Single-source table reading, original board not in hand; "math is exiting" runs against the reality that labs still showcase competition math loudly — a possible reconciliation is that showcasing and training priority are two different things; we keep the tension and don't reconcile it. Judgment update: "K3 can substitute for Opus 4.8" needs a reframe: the two capability profiles genuinely overlap at the per-task level — this is not a case of offsetting strengths that happen to cancel — so stop comparing on a single composite index when selecting models; look at where the per-task residuals fall. For product builders, one operational line: a product that needs strong math cannot expect the next general-purpose flagship to improve it automatically — pull the math item out and measure it separately when you select.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section