SecondSourceJudgment rebuilt from primary sources
Research · Sep 9, 2026

At the checkpoint where one model refused harmful prompts best, it also refused 74% of plainly safe ones.

Model watch · Today (published September 8)

From the Sep 9, 2026 daily brief

The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that across three public safety benchmarks the unsafe-response rate fell from 26.26% to 0.14% — a clean win on that side alone. At the same checkpoint, on XSTest, which exists to measure over-refusal, the false-refusal rate rose from 2.00% to 74.00%. XSTest is a set of prompts that sound dangerous but are harmless. In their own words:

Reported alone, those numbers look like a clean win. They are not. At the same checkpoint, over-refusal on XSTest rises from 2.00% to 74.00%. The configuration with the lowest harmful-response rate is also the one that refuses nearly three quarters of plainly safe prompts. It is a blunt refusal machine, not a safer model, and you cannot see that unless you measure the benign side.

(Multiverse Computing CAI, 2026-09-08). Why topic-level guardrails are not enough, in their own example: the same base model may be deployed as a civic-education tutor or as a public-sector assistant; both should answer factual questions about elections, but only one needs to refuse "write me targeted political manipulation copy." Their fix is to build prompts in pairs that share a topic and differ only in intent. With that data added, false refusals on the should-answer side fell from 32.94% to 4.16%, while the should-refuse side slipped only from 91.88% to 87.72%. When you ask a vendor for safety numbers, ask for the false-refusal rate in the same breath: one side alone cannot show you the blunting. ⚠️ This is a research team introducing its own unreviewed paper — but it volunteered the 74% that works against it, which in our ledger counts in its favor.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section