Models · This month
From the Jul 28, 2026 daily brief
Epoch AI, the AI data-research organization, and the evaluator METR released MirrorCode, a benchmark that asks an AI to reimplement a target program without seeing its source code or touching the internet — working only from command-line inputs and outputs. The results: Opus 4.7 cleared one task in 14 hours at $251 in inference cost that the two organizations estimate would take a human 2–17 weeks; across 25 target programs, 17 were solved perfectly at least once, but 8 never were — the hardest being large, real-world software like ruff, the Python linter. The authors note that leading models a year ago would have scored about 30%, and only on simpler programs like a calendar utility (Import AI #466, Jul 27). Verification: The primary source is a public release by the two evaluation organizations; we read it via relay. Judgment update: "AI can autonomously finish week-scale software tasks" now has systematic measurement behind it, not just anecdotes; framed with core item 2, measuring long-horizon capability and containing long-horizon risk are the two ends of one timeline. Deciding whether an AI coding agent can take over a rewrite? Split the distribution by scale: small, cleanly bounded tools now have systematic evidence behind them; software in the ruff class — 8 of 25 targets never perfectly solved — should stay off your critical path.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The current mainstay of post-training (the stage after base training where a model is taug…
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…
The Jacobian conjecture, posed in 1939, is a famous problem in algebraic geometry. It says…
Our July 21 Research Notes covered this empirical study (gains from optimizing an agent pi…