Model watch
From the Aug 23, 2026 daily brief
— "a cheap open-weight model plus orchestration beats expensive closed-source" is its commercial story, unmediated. There is also a scoring question large enough to flip the ranking: in the first comparison Claude Fable 5 recorded 16 infrastructure errors caused by model routing, and counting them all as failures puts it at 67.3% against the 69.7% officially reported — a 2.4-point spread between two scoring conventions, against a first-attempt gap of only 0.7 points, so changing the convention turns the conclusion over. Costs on the GLM side are given at index level with no per-run detail, so the two sides are not equally auditable. What you can take away, then, is the test — measure how far two models' failures overlap before deciding whether to chain them — and not the conclusion about which one is stronger.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…