August's AI incident wave was mostly backlog; counts track disclosure, not danger — resolves 2026-08-24
Survival test of our call that the August 2026 incident wave was a two-stage chain of disclosures (one new incident plus three batches of older incidents from April to July), and that counting incidents across labs measures the reporting and disclosure infrastructure, not how dangerous the models are.
Sources still being linked up — the node(s) behind this call don't have a clickable original link resolved yet. Logged for follow-up.
Short window: (1) The replies due 2026-08-24 to the congressional oversight letters (17 questions to Anthropic, 23 to OpenAI). Question 1.c.1 asks directly whether the company learned of it only after OpenAI's disclosure, which tests the causal link; also, whether our account of detection on the affected organizations' side is overturned. (2) Week by week through 2026-10-31: if any lab, or Irregular, discloses a new incident of the same type that happened after August 2026 and was not dug up after the fact, the "mostly older incidents" reading narrows (an observation window we set ourselves). (3) The three-way review by METR (Anthropic in dialogue, AISI intending to take part), and a white paper or customer count from Irregular. Long window: whether the markup of the FRONTIER Act (H.R. 9925) adds provisions on evaluation environments or evaluation vendors (trackable in the congress.gov actions); and the first case of the EU AI Act's Article 73 (applicable since 2026-08-02) or California's SB 53 being applied to an incident in an evaluation setting.
Aug 24, 2026
unresolvable (2026-10-04)
Due 2026-08-24. The first short-window item has a reading: in its August 24 reply, Anthropic said it began looking back on July 23, after OpenAI's disclosure on July 21, and did not say that the affected organizations had detected anything, so reversal conditions (2) and (3) were not triggered. But the card was entered without hit conditions or a source of record written down in advance, and the conditions on neither side were triggered, so under our grading scale the result is: cannot be determined. We did not void the card, because that would hide how often our criteria end up undeterminable.