SecondSourceJudgment rebuilt from primary sources
Research · Aug 23, 2026

One vendor ran the same method on two head-to-heads and reached opposite conclusions. What you can actually take away is the test for whether to chain two models together, not who won.

Model watch · This week (posted August 21)

From the Aug 23, 2026 daily brief

Together AI serves inference on open-weight models, and it published two comparisons on the same day using one method and one problem set — 113 software-engineering tasks, four runs each — changing only the opponent. GLM-5.3 is an open-weight model from Zhipu in China. The first (GLM-5.3 against Claude Fable 5): 69.0% against 69.7% on first-attempt pass rate, effectively level, at US$3.99 against US$21.63 per run — a 5.4x price gap. The second (GLM-5.3 against GPT-5.6 Sol): 69.0% against 72.7%, at US$3.99 against US$8.37. The two produce opposite operational advice — one says do not run both, the other says chain them. The two results come from one method pointed at two different opponents, so the opposite advice is what the test found, not a contradiction in it: when two models fail on the same problems (correlation 0.65), combining them only costs more; when they fail on different ones (0.43), the combination buys real complementarity. The cascade in the second piece — run the cheap model first, escalate when the tests reject the answer — reaches 85.9% at US$6.61 per task: 13.2 points better than the expensive model alone, and 21% cheaper.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section