Model watch
From the Aug 9, 2026 daily brief
"Third axis" is the paper's own framing; we do not endorse it. "Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation" (arXiv 2607.27372) self-reports 4.1× compute efficiency and 6.2× sample efficiency. The shape of the numbers former Meta researcher Armen Aghajanyan reported is unusually clean: as the number of candidates generated per step rises, "the objective worked as intended: selected velocity MSE improved 0.0896→0.0763→0.0703 as K rose. But inference MSE (actual rollouts on held-out data) worsened 0.1437→0.1631→0.1835" — one knob, two metrics, opposite directions, both monotonic (the post, August 4, 2026). Held-out data is data the model never saw in training, reserved specifically to test it; improving on the training objective while degrading on held-out data is the textbook shape of the optimized metric coming apart from the property you actually want. ⚠️ That reversal is reported by a single source, not reproduced by us, and we have not confirmed whether his experimental setup is comparable to the paper's — which is exactly the failure mode this item is criticizing, so it is a lead worth chasing, not an established refutation; we have not read the full paper, and the dispute lies entirely in the experiments section. Three questions worth taking away (none of it requires understanding the technical detail): where was this metric measured, on the training distribution or on held-out data? Was anything removed from the denominator? And is this new, or old work under a new name? Efficiency claims are becoming inputs to valuation models and procurement decisions, and these three questions are the shortest path back to something checkable.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The eval house Artificial Analysis runs its AA-Omniscience board on an unusual design: wro…
Teortaxes ran a two-path check on DeepSeek's official disclosures for its late-2024 V3 mod…
First, which yardstick produced "one point apart": the composite intelligence index from t…
Two lines from the K3 launch materials (verbatim, as relayed by Teortaxes): late in develo…