Model watch · Today (event August 12)
From the Aug 14, 2026 daily brief
Stella Biderman is executive director of EleutherAI, the open-source AI research organisation. Mechanistic interpretability is the research line that tries to open the model up and identify specific circuits and mechanisms inside it. Her argument has two levels. The weak version, in the original post: you do not need a mechanistic model to predict the outcome of interventions or to build good theories — "You do not need a mechanistic model to predict the outcome of interventions or even build good theories!" — backed by three historical analogies: that vaccines predate germ theory, that the second law of thermodynamics was identified by someone holding the wrong model of heat, and that quantum mechanics has no standard mechanistic model. The strong version, in her comment: most mechanistic interpretability research is not really proper science (the post). We pulled up the abstract of the paper she linked and compared them. It is arXiv 2606.06533, titled "Position: Don't Just 'Fix it in Post': A Science of AI Must Study Training Dynamics," and she is the first author — so it is not independent corroboration, it is her own argument. The abstract says a science of AI must go beyond fixing things after the fact and study the training dynamics that produce model behaviour, and it explicitly lists mechanistic interpretability among the advances to be examined, rather than excluding it from science (arXiv 2606.06533). The tweet is much stronger than the paper, and the extra stretch has nothing behind it. ⚠️ We read only the abstract, not the full text; this is a position paper, a genre of agenda-setting rather than empirical result, and it must not be cited as "research shows"; and we did not check the three historical analogies one by one. What this means for a decision: a lot of arguments for investing in interpretability rest on an implicit premise — understand the mechanism, and only then can you certify or intervene — and what she is arguing is that the arrow does not hold. That is not the same as saying interpretability is useless. It is saying that if your reason for investing rests on "mechanistic understanding is a necessary precondition," that reason needs replacing. We are not ruling on which side is right — our record holds other material whose core value is a mechanistic finding, neither side has been refuted, so the contradiction stays.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…