← All verdict cards · Verify this card's timestamp ↗
The length of task an AI can finish on its own will pass 8 hours — resolves 2026-11-15
The debate: the evaluation group METR converts how long a task AI can finish on its own into how long a human expert would take, and this line is currently the most accepted yardstick for when AI can take over long pieces of work.
Why it matters: once task length passes a full working day, whole jobs can be handed over instead of fragments, which changes the math of staffing and outsourcing.
Who is on each side: one camp thinks long-running ability is already unlocked; the other thinks the key breakthrough has not arrived yet.
The successor METR benchmark reports an 80%-reliability task time horizon above 8 hours, against a 2026-05 baseline of 3-4 hours.
BG2 Pod(Gerstner)× Atreides(Gavin Baker/Andrew Fox)× Altimeter(Clark Tang)
“there was a Copart tweet about this yesterday. He…”
Dwarkesh Podcast
“if you look at the size of a Tesla, and if you…”
Dwarkesh Podcast (Sholto, then at Google Gemini and later at Anthropic; Trenton, interpretability research at Anthropic)
“I think that's more about nines of reliability and…”
arXiv (Kwa/West/Becker...Barnes/Chan,METR)
“frontier AI time horizon has been doubling…”
Compare METR's next published 80%-reliability horizon against the 8-hour threshold: above is a hit, at or below is a miss; if METR publishes no new reading by the due date, the call is unresolvable.
What would overturn it
A reading that stalls in the 3-4 hour range or moves backwards supports the view that the unlock for long-running work has not arrived yet.
Nov 15, 2026
2026-11-15