Model watch · Trend watch (original paper June 2025 / follow-up paper 2026)
From the Sep 4, 2026 daily brief
CyberGym is a public benchmark from a Berkeley team, with 1,507 real vulnerabilities across 188 open-source projects. The data comes from OSS-Fuzz, the open-source fuzzing platform Google runs, so the vulnerabilities are by construction already known and on record. The task, as the paper defines it, is for the model to produce a proof-of-concept test that reproduces the vulnerability, and all the model gets is a text description of the vulnerability and the matching codebase. The paper's own best combination in 2025 reached roughly a 20% success rate. Within a year, two frontier labs' self-reported scores jumped to 85.6 and 86.2, and neither has said which difficulty band or subset it tested, and no third party has reproduced either score (CyberGym, 2025-06). The most informative line is the opening of the same authors' follow-up paper, which concedes that existing security evaluations of AI systems fail to cover the end-to-end lifecycle of real-world vulnerability discovery and repair (CyberGym-E2E, 2026). One thing to take away and use: when you see any security benchmark score, first ask what the test gives the model and what it asks the model to do.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…