Model watch · This week (published September 1)
From the Sep 2, 2026 daily brief
The Allen Institute for AI published BenchMIRT on September 1. It uses a multidimensional form of item response theory (the method psychological testing uses to estimate ability from answer patterns) to take a benchmark apart question by question and ask what it actually measures. Fed 100 models × 16 benchmarks × more than 34,000 questions, and told nothing about what any benchmark was supposed to measure, it produced two axes of its own: safety and general reasoning. Here is the counterintuitive part. BBQ, a social-bias benchmark, and WMDP, which tests dual-use hazardous knowledge, both sit under safety by convention, and both correlate far more strongly with general reasoning. Stronger reasoning even means a lower WMDP score, because that benchmark counts refusing to answer as correct (Hugging Face blog, 09-01). Read alongside item 2 of today's main line: Korea's "benchmarks are worth 40 points" means four-tenths of the score comes from a set of questions whose own subject matter nobody has taken apart yet. If you are picking a model on safety-benchmark scores, first check whether the benchmark counts a refusal as a right answer — otherwise what you select for is the model best at avoiding the question, not the safest one. ⚠️ Every model used to train it was released before March 2025, so its behavior on the current generation is unknown.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…