SecondSourceJudgment rebuilt from primary sources
Research · Sep 4, 2026

The ruler all three labs cite measures reproducing a vulnerability you have already been told about.

Model watch · Trend watch (original paper June 2025 / follow-up paper 2026)

From the Sep 4, 2026 daily brief

CyberGym is a public benchmark from a Berkeley team, with 1,507 real vulnerabilities across 188 open-source projects. The data comes from OSS-Fuzz, the open-source fuzzing platform Google runs, so the vulnerabilities are by construction already known and on record. The task, as the paper defines it, is for the model to produce a proof-of-concept test that reproduces the vulnerability, and all the model gets is a text description of the vulnerability and the matching codebase. The paper's own best combination in 2025 reached roughly a 20% success rate. Within a year, two frontier labs' self-reported scores jumped to 85.6 and 86.2, and neither has said which difficulty band or subset it tested, and no third party has reproduced either score (CyberGym, 2025-06). The most informative line is the opening of the same authors' follow-up paper, which concedes that existing security evaluations of AI systems fail to cover the end-to-end lifecycle of real-world vulnerability discovery and repair (CyberGym-E2E, 2026). One thing to take away and use: when you see any security benchmark score, first ask what the test gives the model and what it asks the model to do.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section