Model watch · Evals Trend watch
From the Aug 5, 2026 daily brief
The eval house Artificial Analysis runs its AA-Omniscience board on an unusual design: wrong answers cost points, admitting you don't know costs nothing — so it measures knowledge breadth and the tendency to refuse to fabricate at the same time. Teortaxes (the same commentator as today's item #1) read the new round for something other than rankings: the model at the very top on knowledge is also among the most prone to making things up (Teortaxes, July 21). That independently reproduces Artificial Analysis's own methodology finding from January: hallucination rate doesn't correlate with overall intelligence, while knowledge accuracy correlates strongly with total parameter count and barely at all with active parameters (Latent Space interview archive, January 2026) — and the same regularity reproducing across an entirely new crop of models upgrades a one-time snapshot into a pattern with a time axis. The selection read: for cold knowledge and long-tail facts, go toward big total parameter counts; for "would rather say I don't know" reliability, look at the post-training recipe and refusal behavior — two separate ledgers, and hiring the knowledge champion as your automatic judge means buying its fabrication habit too. Their read-out carried no scores; treat the citation as direction.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
Teortaxes ran a two-path check on DeepSeek's official disclosures for its late-2024 V3 mod…
First, which yardstick produced "one point apart": the composite intelligence index from t…
Two lines from the K3 launch materials (verbatim, as relayed by Teortaxes): late in develo…
Overnight's 15 new papers are still unprocessed and this column has no single-day incremen…