SecondSourceJudgment rebuilt from primary sources
Research · Sep 2, 2026

Somebody took psychometrics to the benchmarks and found half of "safety" is measuring reasoning.

Model watch · This week (published September 1)

From the Sep 2, 2026 daily brief

The Allen Institute for AI published BenchMIRT on September 1. It uses a multidimensional form of item response theory (the method psychological testing uses to estimate ability from answer patterns) to take a benchmark apart question by question and ask what it actually measures. Fed 100 models × 16 benchmarks × more than 34,000 questions, and told nothing about what any benchmark was supposed to measure, it produced two axes of its own: safety and general reasoning. Here is the counterintuitive part. BBQ, a social-bias benchmark, and WMDP, which tests dual-use hazardous knowledge, both sit under safety by convention, and both correlate far more strongly with general reasoning. Stronger reasoning even means a lower WMDP score, because that benchmark counts refusing to answer as correct (Hugging Face blog, 09-01). Read alongside item 2 of today's main line: Korea's "benchmarks are worth 40 points" means four-tenths of the score comes from a set of questions whose own subject matter nobody has taken apart yet. If you are picking a model on safety-benchmark scores, first check whether the benchmark counts a refusal as a right answer — otherwise what you select for is the model best at avoiding the question, not the safest one. ⚠️ Every model used to train it was released before March 2025, so its behavior on the current generation is unknown.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section