SecondSourceJudgment rebuilt from primary sources
Research · Aug 22, 2026

Two things are new today.

Model watch

From the Aug 22, 2026 daily brief

The first is coming back empty-handed. We ran a targeted search for independent replication or refutation, and all four hits trace to the same origin: the paper itself, the research group's project page, OpenAI's own announcement, and the study's dedicated site. After de-duplication that is still one source, so our internal confidence score stays at 0.55 (out of 1, meaning the evidence does not yet carry either side; how we score). Nearly a year after publication with no outside replication is itself worth writing down: the central negative finding about safety-training effectiveness has so far been measured only by the people who proposed it. The second is a method detail. That causal evidence works by editing the model's reasoning text directly: "add mentions of evaluation awareness, models take fewer covert actions; removing those mentions makes covert actions more likely" (the study's FAQ page). This points at a self-referential limit: the experiment used to show that evaluation awareness contaminates the results depends on exactly the human-readable reasoning that the same paper warns is degrading. If models stop reasoning in standard English, the causal experiment and the warning it supports both stop working, with no way left to check either. Secondhand summaries carry several finer ablation figures, but we could find no matching text in either the paper's abstract or the official FAQ, so by our own rules we neither use them nor relay them.

How to use it: whenever a lab cites "safety training cut misbehaviour by X%," this is the standard piece to put on the table. Ask two things: has a second team reproduced that drop in its own environment, and is the reasoning text the measurement depends on still readable on your models.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section