SecondSourceJudgment rebuilt from primary sources
Research · Sep 20, 2026

Cross-training each year's open-source model designs and training recipes from 2019 to 2025 against each year's corpora puts the data side's compute-equivalent gain at 3.24 times the model side's — a ratio of two gain multiples, not a share of contribution. The same day, the result was relayed as "model size scaling is dead."

Model watch · This month (originally posted September 8)

From the Sep 20, 2026 daily brief

The experiment takes representative open-source model designs and training recipes from each of the six years between 2019 and 2025, crosses them with each year's training corpus, and retrains at a range of small scales. The result: a compute-equivalent gain of 12.0x on the data side and 3.7x on the model side; divide one by the other and you get 3.24. That is a ratio of two gain multiples, not a share of contribution, and the author says the two gains stack almost independently, with good data helping every architecture about equally (Dwarkesh Patel, 2026-09-08). Compute-equivalent gain means: to reach the same progress without this improvement, how many times more compute would you have to spend. ⚠️ Three reservations travel together. The scale range is described in the original only as "a range of small scales." Pretraining is the stage where a base model is trained on a large body of text first, and conclusions from experiments like this are extremely sensitive to scale. Which year's recipe counts as "representative" is a choice that itself steers the result. And there are no error bars, no peer review and no independent replication. The same day, Sara Hooker, CEO of Adaption Labs, cited the experiment, and the version she stated was "Pretraining model size scaling is dead because transformers are saturated" and "Only gains now are data" (Sara Hooker, 2026-09-08). ⚠️ Her company's whole argument is that learning from experience and data beats scaling models up, so that line can't be cited as a neutral technical observation. The real use of this item is the reminder to go back and check the numbers: the original experiment measured 3.7x still sitting on the model side, and the relayed version became "only data is left." When you meet a sentence of the form "X is dead," the first move is to go back to the number it rests on and see whether that number actually says zero.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section