SecondSourceJudgment rebuilt from primary sources
Research · Aug 13, 2026

The paper being circulated as proof that Chinese models were distilled has two channels whose numbers differ by a factor of a hundred thousand — and what went viral is the weak one.

Model watch · Today (a dispute about measurement, events August 11 to 12)

From the Aug 13, 2026 daily brief

Distillation means using one model's output (the teacher) to train another (the student). A paper has been widely cited in recent days as proof that Chinese open-weight models were distilled from Anthropic's Opus. @bookwormengr — an X account whose institutional affiliation we could not verify, and a declared opponent of the distillation claim — walked through the paper paragraph by paragraph and gave these numbers: to make a model reproduce sixteen consecutive tokens of Opus's reasoning verbatim, the theoretical number of independent attempts required is around 10^10 for the most extractable model in the set, with the rest landing between 10^14 and 10^16; but the same thing on the visible answer takes only about 100,000. Two channels, roughly a hundred thousand times apart, and the one that went viral as evidence is the latter. He also quotes the paper against itself: "The paper states directly that these results do not support verbatim memorisation of the reasoning" (the post). ⚠️ The biggest problem has to go in the sentence itself: we do not have the paper. We know neither authors nor identifier nor title; every number reaches us through a relay with a declared position, and the attempt counts are theoretical estimates rather than measurements actually run. So the conclusion can only be written as "under the relayed numbers, this widely circulated evidence does not hold up the conclusion drawn from it" — it cannot be written as "Chinese models were not distilled." What you can take away directly is three questions to ask. Next time you see "model X has been proven to be a distillation," ask: does the evidence sit in the reasoning channel or the answer channel? Is the test set a well-known public benchmark? Was a within-family baseline run as a control? The reason for the second question is that if the test set is public, "that question was already in the training data" is an equally compatible and far simpler explanation. And one counterintuitive note: in that paper's experiments, the model behind the flood of "I'm Claude" screenshots online turns out to be clean at the token level. Getting your own identity wrong and having your weights copied are two things that can happen separately.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section