SecondSourceThe Second Source on AI

← All verdict cards · Verify this card's timestamp ↗

The successor METR benchmark reports… — resolves 2026-11-15

Pending

The successor METR benchmark reports… — resolves 2026-11-15

Trend call · Asserted on Jul 31, 2026

The call

The successor METR benchmark reports an 80%-reliability task time horizon above 8 hours, against a 2026-05 baseline of 3-4 hours.

Sources

arXiv (Kwa/West/Becker...Barnes/Chan,METR)
「frontier AI time horizon has been doubling approxi…」

BG2 Pod(Gerstner)× Atreides(Gavin Baker/Andrew Fox)× Altimeter(Clark Tang)
「there was a Copart tweet about this yesterday. He …」

Dwarkesh Podcast
「if you look at the size of a Tesla, and if you loo…」

Dwarkesh Podcast(Sholto 時任 Google Gemini,後轉 Anthropic;Trenton 任 Anthropic interp)
「I think that's more about nines of reliability and…」

The receipt

Compare METR's next published 80%-reliability horizon against the 8-hour threshold: above is a hit, at or below is a miss; if METR publishes no new reading by the due date, the call is unresolvable.

Resolves

Nov 15, 2026

2026-11-15

Outcome

Not due yet — no public outcome to report.

What we learned

Not resolved yet — here is the condition we pre-committed to that would overturn this call (check it yourself, you don't have to wait):

A reading that stalls in the 3-4 hour range or moves backwards supports the view that the unlock for long-running work has not arrived yet.

The successor METR benchmark reports… — resolves 2026-11-15