← All verdict cards · Verify this card's timestamp ↗
The successor METR benchmark reports… — resolves 2026-11-15
The successor METR benchmark reports an 80%-reliability task time horizon above 8 hours, against a 2026-05 baseline of 3-4 hours.
arXiv (Kwa/West/Becker...Barnes/Chan,METR)
「frontier AI time horizon has been doubling approxi…」
BG2 Pod(Gerstner)× Atreides(Gavin Baker/Andrew Fox)× Altimeter(Clark Tang)
「there was a Copart tweet about this yesterday. He …」
Dwarkesh Podcast
「if you look at the size of a Tesla, and if you loo…」
Dwarkesh Podcast(Sholto 時任 Google Gemini,後轉 Anthropic;Trenton 任 Anthropic interp)
「I think that's more about nines of reliability and…」
Compare METR's next published 80%-reliability horizon against the 8-hour threshold: above is a hit, at or below is a miss; if METR publishes no new reading by the due date, the call is unresolvable.
Nov 15, 2026
2026-11-15
Not due yet — no public outcome to report.
Not resolved yet — here is the condition we pre-committed to that would overturn this call (check it yourself, you don't have to wait):
A reading that stalls in the 3-4 hour range or moves backwards supports the view that the unlock for long-running work has not arrived yet.
