SecondSourceJudgment rebuilt from primary sources
Research · Sep 4, 2026

(original paper December 20, 2025 / countervailing evidence September 3) A load-bearing finding in an OpenAI paper was offset by the same company's system card eight and a half months later.

Model watch · Evidence update

From the Sep 4, 2026 daily brief

On December 20, 2025, OpenAI published "Monitoring Monitorability", which turned how well a chain of thought can be monitored into a measurable metric. One of its key findings, word for word: "RL optimization does not materially decrease monitorability even at the current frontier scale" (arXiv 2512.18311). The Astra system card of September 3, 2026 says "CoT controllability for Astra especially increases over the course of RL training". Rising controllability is bad news, for the reasons given above. After rechecking, our handling is to lower the credibility of the older finding without marking it dead. The cause is not settled. If the capability jump is the main driver, the paper was not wrong within the scope it claimed, and it wrote "current scale" into that scope itself as the limit of extrapolation. The lesson to record is not "don't trust vendor self-reports"; our process already guards against that. The lesson is that when a conclusion carries an explicit scope, the scope is the first thing to get dropped when someone else cites it.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section