SecondSourceJudgment rebuilt from primary sources
Research · Sep 5, 2026

(original posts August 14, 2026 / published version August 21; settled today) The security evaluation we held for three weeks turned its prettiest sentence around once it was published in full.

Model watch · Evidence update

From the Sep 5, 2026 daily brief

On August 14 we recorded a post: an evaluator with early access to Z.ai's open-weight flagship, GLM-5.3, measured how well it finds vulnerabilities and produced four flattering numbers. We did not take them at the time, and we marked three gaps. None of the four numbers had a reproducible denominator. There was no statement of whether the vulnerabilities became public after the training cutoff. And we had not read the last two posts in the thread. The second gap mattered most: if the vulnerabilities were already public before the cutoff, the model may simply have memorized the answers. Checking today, we found that the same people, the security firm Aikido Security, had published the full methodology a week later (Aikido Security, 2026-08-21). All three gaps now have answers, and the net change after settling splits two ways rather than mapping onto those three gaps one by one. In our favor: the denominator is 32 recently disclosed vulnerabilities, with each model run through all 32 cases three times, 96 runs in total, and the authors write that "We were careful about dataset freshness" — the selection was built to cut the chance that the models had already seen those vulnerabilities. The decisive worry we flagged had occurred to them as well. Two points cut against the original post. Two terms first: recall is the share of the vulnerabilities that should have been found that actually were, and precision is the share of what gets reported that is real, with the rest being false positives. Recall moved from 24/32 to 25/32. The cost comparison went from "40% cheaper than GPT-5.6-Terra" to "65.5% cheaper than Sol and 69.8% cheaper than Opus 5", changing both the comparison target and the margin. And the prettiest line in the original post — that it produced fewer false positives than other high-recall models — is overturned in the published version by its own authors: "The open models caught up with the frontier at much cheaper rates but also produced the most false leads for the pipeline to reject." The open models generate the most false positives. ⚠️ A post upgraded into a blog post does not become two independent sources. This remains one team's own evaluation, and our confidence stays capped at the single-source ceiling. ⚠️ The vulnerability list and its date distribution are still unpublished, so a reader cannot check for themselves how fresh "fresh" is. You can buy recall. You cannot buy precision. Ask vendors to report a false-positive rate alongside recall; a quote that gives you recall alone is not comparable to anything. This item also pushes back on item 2 of today's main line: if the evaluators are right, holding the open weights is not the same as holding usable attack capability, and a framework nobody has measured sits in between. If that holds, all three readings above point to expiry dates that need to move later.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section