Model watch · This week (posted August 21)
From the Aug 23, 2026 daily brief
Together AI serves inference on open-weight models, and it published two comparisons on the same day using one method and one problem set — 113 software-engineering tasks, four runs each — changing only the opponent. GLM-5.3 is an open-weight model from Zhipu in China. The first (GLM-5.3 against Claude Fable 5): 69.0% against 69.7% on first-attempt pass rate, effectively level, at US$3.99 against US$21.63 per run — a 5.4x price gap. The second (GLM-5.3 against GPT-5.6 Sol): 69.0% against 72.7%, at US$3.99 against US$8.37. The two produce opposite operational advice — one says do not run both, the other says chain them. The two results come from one method pointed at two different opponents, so the opposite advice is what the test found, not a contradiction in it: when two models fail on the same problems (correlation 0.65), combining them only costs more; when they fail on different ones (0.43), the combination buys real complementarity. The cascade in the second piece — run the cheap model first, escalate when the tests reject the answer — reaches 85.9% at US$6.61 per task: 13.2 points better than the expensive model alone, and 21% cheaper.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…