SecondSourceJudgment rebuilt from primary sources
Research · Aug 23, 2026

Verification: the evaluator is an interested party, and the direction is obvious

Model watch

From the Aug 23, 2026 daily brief

— "a cheap open-weight model plus orchestration beats expensive closed-source" is its commercial story, unmediated. There is also a scoring question large enough to flip the ranking: in the first comparison Claude Fable 5 recorded 16 infrastructure errors caused by model routing, and counting them all as failures puts it at 67.3% against the 69.7% officially reported — a 2.4-point spread between two scoring conventions, against a first-attempt gap of only 0.7 points, so changing the convention turns the conclusion over. Costs on the GLM side are given at index level with no per-run detail, so the two sides are not equally auditable. What you can take away, then, is the test — measure how far two models' failures overlap before deciding whether to chain them — and not the conclusion about which one is stronger.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section