Model watch · Today (posted September 10)
From the Sep 10, 2026 daily brief
The instrument here belongs to someone else — an evaluation called CoT-Control, which appends instructions about the reasoning trace to a question ("think only in lowercase," "do not use this word") and grades whether the trace complied. The results cluster by lab across three downloadable-weight model families from three different labs: the GLM-5 series from Z.AI (formerly Zhipu AI, a Chinese frontier lab) lands between 6% and 13%, one DeepSeek model at 11%, and every Qwen 3.5 checkpoint — Qwen is Alibaba's open-weight family — from 9B to 397B sits at zero (Alignment Forum, 2026-09-10). ⚠️ Only open-weight models can be measured this way, because only they expose the real reasoning trace — closed models hand you a summary — so this reading cannot reach the model in today's main line. It has no peer review and no named author. Why it is useful: it reclassifies controllability from an inevitable by-product of growing capability into an engineering choice, and engineering choices are the sort of thing procurement contracts and regulations can specify.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…
Six academic papers reached the reading list last night and none was finished today. Three…