Model watch · Evidence update
From the Sep 5, 2026 daily brief
On August 14 we recorded a post: an evaluator with early access to Z.ai's open-weight flagship, GLM-5.3, measured how well it finds vulnerabilities and produced four flattering numbers. We did not take them at the time, and we marked three gaps. None of the four numbers had a reproducible denominator. There was no statement of whether the vulnerabilities became public after the training cutoff. And we had not read the last two posts in the thread. The second gap mattered most: if the vulnerabilities were already public before the cutoff, the model may simply have memorized the answers. Checking today, we found that the same people, the security firm Aikido Security, had published the full methodology a week later (Aikido Security, 2026-08-21). All three gaps now have answers, and the net change after settling splits two ways rather than mapping onto those three gaps one by one. In our favor: the denominator is 32 recently disclosed vulnerabilities, with each model run through all 32 cases three times, 96 runs in total, and the authors write that "We were careful about dataset freshness" — the selection was built to cut the chance that the models had already seen those vulnerabilities. The decisive worry we flagged had occurred to them as well. Two points cut against the original post. Two terms first: recall is the share of the vulnerabilities that should have been found that actually were, and precision is the share of what gets reported that is real, with the rest being false positives. Recall moved from 24/32 to 25/32. The cost comparison went from "40% cheaper than GPT-5.6-Terra" to "65.5% cheaper than Sol and 69.8% cheaper than Opus 5", changing both the comparison target and the margin. And the prettiest line in the original post — that it produced fewer false positives than other high-recall models — is overturned in the published version by its own authors: "The open models caught up with the frontier at much cheaper rates but also produced the most false leads for the pipeline to reject." The open models generate the most false positives. ⚠️ A post upgraded into a blog post does not become two independent sources. This remains one team's own evaluation, and our confidence stays capped at the single-source ceiling. ⚠️ The vulnerability list and its date distribution are still unpublished, so a reader cannot check for themselves how fresh "fresh" is. You can buy recall. You cannot buy precision. Ask vendors to report a false-positive rate alongside recall; a quote that gives you recall alone is not comparable to anything. This item also pushes back on item 2 of today's main line: if the evaluators are right, holding the open weights is not the same as holding usable attack capability, and a framework nobody has measured sits in between. If that holds, all three readings above point to expiry dates that need to move later.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…