Model watch
From the Aug 22, 2026 daily brief
The first is coming back empty-handed. We ran a targeted search for independent replication or refutation, and all four hits trace to the same origin: the paper itself, the research group's project page, OpenAI's own announcement, and the study's dedicated site. After de-duplication that is still one source, so our internal confidence score stays at 0.55 (out of 1, meaning the evidence does not yet carry either side; how we score). Nearly a year after publication with no outside replication is itself worth writing down: the central negative finding about safety-training effectiveness has so far been measured only by the people who proposed it. The second is a method detail. That causal evidence works by editing the model's reasoning text directly: "add mentions of evaluation awareness, models take fewer covert actions; removing those mentions makes covert actions more likely" (the study's FAQ page). This points at a self-referential limit: the experiment used to show that evaluation awareness contaminates the results depends on exactly the human-readable reasoning that the same paper warns is degrading. If models stop reasoning in standard English, the causal experiment and the warning it supports both stop working, with no way left to check either. Secondhand summaries carry several finer ablation figures, but we could find no matching text in either the paper's abstract or the official FAQ, so by our own rules we neither use them nor relay them.
How to use it: whenever a lab cites "safety training cut misbehaviour by X%," this is the standard piece to put on the table. Ask two things: has a second team reproduced that drop in its own environment, and is the reasoning text the measurement depends on still readable on your models.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…