SecondSourceJudgment rebuilt from primary sources
Research · Jul 28, 2026

(the primary release carries no explicit date — confirmable only to July 2026; picked up July 27 via commentary) New benchmark MirrorCode: an AI spends 14 hours and $251 on a software-rewrite task humans would need 2–17 weeks for

Models · This month

From the Jul 28, 2026 daily brief

Epoch AI, the AI data-research organization, and the evaluator METR released MirrorCode, a benchmark that asks an AI to reimplement a target program without seeing its source code or touching the internet — working only from command-line inputs and outputs. The results: Opus 4.7 cleared one task in 14 hours at $251 in inference cost that the two organizations estimate would take a human 2–17 weeks; across 25 target programs, 17 were solved perfectly at least once, but 8 never were — the hardest being large, real-world software like ruff, the Python linter. The authors note that leading models a year ago would have scored about 30%, and only on simpler programs like a calendar utility (Import AI #466, Jul 27). Verification: The primary source is a public release by the two evaluation organizations; we read it via relay. Judgment update: "AI can autonomously finish week-scale software tasks" now has systematic measurement behind it, not just anecdotes; framed with core item 2, measuring long-horizon capability and containing long-horizon risk are the two ends of one timeline. Deciding whether an AI coding agent can take over a rewrite? Split the distribution by scale: small, cleanly bounded tools now have systematic evidence behind them; software in the ruff class — 8 of 25 targets never perfectly solved — should stay off your critical path.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section