SecondSourceJudgment rebuilt from primary sources
Research · Aug 14, 2026

The paper a senior researcher linked to back her line that "most mech interp research isn't really proper science" turns out to be her own, and its abstract does not say that.

Model watch · Today (event August 12)

From the Aug 14, 2026 daily brief

Stella Biderman is executive director of EleutherAI, the open-source AI research organisation. Mechanistic interpretability is the research line that tries to open the model up and identify specific circuits and mechanisms inside it. Her argument has two levels. The weak version, in the original post: you do not need a mechanistic model to predict the outcome of interventions or to build good theories — "You do not need a mechanistic model to predict the outcome of interventions or even build good theories!" — backed by three historical analogies: that vaccines predate germ theory, that the second law of thermodynamics was identified by someone holding the wrong model of heat, and that quantum mechanics has no standard mechanistic model. The strong version, in her comment: most mechanistic interpretability research is not really proper science (the post). We pulled up the abstract of the paper she linked and compared them. It is arXiv 2606.06533, titled "Position: Don't Just 'Fix it in Post': A Science of AI Must Study Training Dynamics," and she is the first author — so it is not independent corroboration, it is her own argument. The abstract says a science of AI must go beyond fixing things after the fact and study the training dynamics that produce model behaviour, and it explicitly lists mechanistic interpretability among the advances to be examined, rather than excluding it from science (arXiv 2606.06533). The tweet is much stronger than the paper, and the extra stretch has nothing behind it. ⚠️ We read only the abstract, not the full text; this is a position paper, a genre of agenda-setting rather than empirical result, and it must not be cited as "research shows"; and we did not check the three historical analogies one by one. What this means for a decision: a lot of arguments for investing in interpretability rest on an implicit premise — understand the mechanism, and only then can you certify or intervene — and what she is arguing is that the arrow does not hold. That is not the same as saying interpretability is useless. It is saying that if your reason for investing rests on "mechanistic understanding is a necessary precondition," that reason needs replacing. We are not ruling on which side is right — our record holds other material whose core value is a mechanistic finding, neither side has been refuted, so the contradiction stays.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section