Every factual error we shipped to readers and later had to fix is listed here — with which layer failed, and what we changed because of it. There is no filter on this list.
-
Corrected 2026-10-06Chinese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- The 7/25 deep column (No. 14) said "only 11% of teams switched model vendors in the past year; 66% upgraded to newer models inside their existing vendor," linking to Menlo Ventures' year-end 2025 report, and said Anthropic's inference-serving gross margin, "per internal and analyst accounts," climbed "from 38% in 2024" to around 65% in 2026. The deep-dive section of the 7/25 brief and the 8/4 brief linked the same 11% to the same year-end report.
- What was actually the case
- In the year-end report, 11% is the market share of open-source models; the switching figure comes from Menlo's 2025 mid-year report: "Only 11% of teams report changing model providers in the past year. But 66% upgraded to newer models from their existing vendor." The margin figures are an estimate by the research firm SemiAnalysis: "inference gross margins are now in the mid 60s, up from 38% in 2025 and -94% in 2024" — 38% is for 2025, not 2024.
- Which layer failed
- Bad data — the numbers were right, the source and year were not: the 11% was attached to the URL of a different report from the same firm, the margin year was copied one year off, and an analyst firm's estimate was described as "internal and analyst accounts." The trend tree in the same piece says -94% for 2024, and the draft was not checked against it.
- What we did about it
- On 2026-10-06 the archive page moved the 11% link to Menlo's 2025 mid-year report, changed the margin passage to "per an estimate by the research firm SemiAnalysis, from 38% in 2025" with the SemiAnalysis source cited, and added a note at the top; the Chinese edition was corrected the same day with its own note. The same link in the deep-dive section of the 7/25 brief (all three languages) and in the 8/4 brief (all three languages) was corrected the same day with notes. The conclusion does not rest on either passage and is unchanged. The emails sent on 7/25 and 8/4 are the originals.
Editions affected: /deep/2026-07-25/, /en/deep/2026-07-25/, 2026-07-25, 2026-08-04
-
Corrected 2026-10-05Chinese editionBad inference — the inputs were right, the reasoning or the wording was not
- What we wrote
- The verdict card on whether large models will become swappable commodities (call made July 25) had a Chinese headline saying “all three numbers currently point the other way,” and its background said we “checked three real numbers and take the opposite side”; the English and Japanese cards said the thesis fails a stress test against all three of its own watchpoints, while the badge on the page says the call is pending.
- What was actually the case
- Only one of the three watchpoints has a before-and-after comparison: Menlo Ventures measured Anthropic rising from 32% in mid-2025 to 40% in December, with the top three at 88% combined. The 11% switching figure comes from a single Menlo survey in mid-2025, with no earlier figure to compare against. The margin item gave no number on the page; what exists is an outside estimate by the research firm SemiAnalysis (inference gross margin rising from 38% in 2025 to the mid-60s in the first quarter of 2026), unaudited. The card does not resolve until July 25, 2027.
- Which layer failed
- Bad inference — the material was right: the July 25 deep column said each of the three readings “comes with a caveat that cannot be ignored,” that the 11% has no reliable baseline, and that the margin figures are unaudited. Rewriting it into the card's plain-language layer turned three watchpoints into “three numbers” and dropped the caveats, and the English and Japanese cards stated our call as an outcome that had already happened.
- What we did about it
- On 2026-10-05 the card pages rewrote the headline and the call in all three languages: the call is stated as ours and resolves on July 25, 2027; each of the three items now gives its number, its baseline and what kind of source it is; and the call section carries a correction note. The last line of the Chinese background was changed to match. The same day the card added the readings from the September 30 interim check: Menlo has not published a new report, and Claude Sonnet 5's introductory price became its standard price, so under item (3) of the receipt the price item was revisited. The direction of the call (only a layered version holds, not across-the-board commoditization) is unchanged.
Editions affected: /cards/v-2026-07-25-01/, /en/cards/v-2026-07-25-01/, /ja/cards/v-2026-07-25-01/
-
Corrected 2026-10-05Chinese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- The 7/28 headline, the first At a glance item and core item 1 said that of ten Chinese labs exactly one, MiniMax, had "followed" K3 with a gate of its own, and that MiniMax's license was "published July 23"; the third At a glance item and the headline of core item 3 said the "pick a model for me" distribution layer "has been shown to make money."
- What was actually the case
- Hugging Face's file history shows the MiniMax M3 license was uploaded with the model page's first commit on June 12 and has not changed since; July 23 is only when the page's description was last updated. It came more than a month before K3's license (July 27), so MiniMax did not follow K3. The only evidence for the distribution-layer item is an OpenRouter board member's own account, with no figures, and the item itself says confidence does not move.
- Which layer failed
- Bad data — the license date was taken from the model page's last-updated date rather than the date the license file was uploaded, and "followed" was written from that wrong date; the same day's deep column already said MiniMax came before K3, and the brief was not checked against it. The distribution-layer headline said more than its own item.
- What we did about it
- On 2026-10-05 the archive page changed the three passages to say that besides K3 only one lab has a similar gate and that it came first, gave the June date, changed the two distribution-layer headlines to "is said to make money," and added a note at the top; the finding that only two of the ten put a revenue gate in their license stands. The Chinese and Japanese editions were corrected the same day with their own notes, and the Japanese deep column of the same day, which put MiniMax's license at July 23 and four days ahead of K3, was corrected with a note. The email sent on 7/28 is the original.
Editions affected: 2026-07-28, /ja/deep/2026-07-28/
-
Corrected 2026-10-05Chinese editionBad inference — the inputs were right, the reasoning or the wording was not
- What we wrote
- The 8/29 deep column (No. 26) said the numerator and denominator of the "two-to-five-times gap" (a trading firm capturing US$200M–500M per MW, and a model company earning US$50M per MW) came from two different companies, that the trading firm was not that model company's customer, citing the transcript; its headline and summary said we had made the judgment "yesterday" in our August 28 issue, and main-line item 1 of the 8/29 issue said our August 28 issue argued it.
- What was actually the case
- The half-sentence the quotation dropped with an ellipsis reads "or Jane Street where they're one of Anthropic's biggest customers": the speaker calls the firm both a contract customer of OpenAI's fast mode and one of Anthropic's biggest customers. US$200M–500M divided by the US$50M on the page is four to ten times; "two to five times" divides by another figure the speaker gave late in the episode, US$100M per MW. The judgment was logged in our ledger on August 28, and that day's public issue did not carry it.
- Which layer failed
- Bad inference — the material was right: the transcript gives both denominators and calls the firm a customer of both labs. The draft took only the US$50M from the opening, the ellipsis happened to remove the half-sentence that contradicted the paragraph's point, and "we said yesterday" was written from the ledger's entry date without checking that day's public issue.
- What we did about it
- On 2026-10-05 the archive page rewrote section 1, the summary, the route paragraph, the section 2 line on the current ceiling and falsification condition 5 to follow the transcript, changed the headline and summary to "the judgment we logged on August 28" with a note that the issue did not carry it, and added a note at the top; the conclusion (the leading indicator cannot discriminate, watch list prices and quotas) does not rest on these and is unchanged. The Japanese edition was corrected the same day and its headline now separates the circulating consolation line from our judgment; main-line item 1 of the 8/29 issue was corrected in all three languages with its own notes. The email sent on 8/29 is the original.
Editions affected: /deep/2026-08-29/, /ja/deep/2026-08-29/, 2026-08-29
-
Corrected 2026-10-04Chinese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- The 'What happens when we're wrong' section of our methodology page said we had 49 calls tracked and 20 public as verdict cards, that nothing had come due yet, and that the table would start keeping score in September; the 'What we haven't done yet' section on the same page said the scoreboard hadn't settled a single call — zero track record.
- What was actually the case
- The ledger kept growing after that, so the count had long stopped being 49; the first call came due on 2026-08-24, and by 10-04 several had come due, each with its result on its own card. The admission-log and live-call counts in the gates section of the same page were also hard-coded at the time and had not been updated since.
- Which layer failed
- Bad data — the ledger figures and the due status were text hard-coded into the site's code in early August, and nothing updated them as the ledger grew and calls came due.
- What we did about it
- From 2026-10-04 the due-status text on the methodology page is generated from the call ledger on every site build: it says only whether calls have come due and where to read the results, with no counts; the card-link blurb was rewritten to match. The 'What we haven't done yet' item now says there is no public overall score yet and why; the gates section drops the hard-coded admission-log count, live-call count and top confidence score, and its line that more items get blocked or parked than admitted, which holds or not depending on how decisions are classified, now says instead that blocked and parked items are logged the same way and can be traced. The Chinese and Japanese pages were corrected the same day.
Editions affected: /methodology/, /en/methodology/, /ja/methodology/
-
Corrected 2026-09-17Chinese editionBad inference — the inputs were right, the reasoning or the wording was not
- What we wrote
- The Chinese edition of 9/14 said the load-bearing sources of 'five main-line items' traced back to six unrelated origins.
- What was actually the case
- The main line on the page has four items; the fifth was moved to the not-yet-verified section, and its source, OpenAI's official announcement, was one of the six origins and is not used by the four remaining items, so it is four items and five origins.
- Which layer failed
- Bad inference — when items were moved from the main line to the not-yet-verified section for length, only the items themselves moved; the inventory sentence built from the main-line list was not recomputed and kept its pre-move count.
- What we did about it
- On 2026-09-17 the archive page was changed to four items and five origins with a note at the top; 'six origins' in the English and Japanese editions was changed to five the same day with their own notes. The email sent on 9/14 is the original.
Editions affected: 2026-09-14
-
Corrected 2026-09-17Chinese editionBad inference — the inputs were right, the reasoning or the wording was not
- What we wrote
- The Chinese edition of 9/13 said the load-bearing sources of 'four main-line items' spanned five kinds, the first being 'official model documentation and a pricing page'.
- What was actually the case
- The main line on the page has two items; the other two were moved to the not-yet-verified section, and the pricing-page source belonged only to a moved item.
- Which layer failed
- Bad inference — when items were moved from the main line to the not-yet-verified section for length, only the items themselves moved; the inventory sentence built from the main-line list was not recomputed and kept its pre-move count.
- What we did about it
- On 2026-09-17 the archive page was changed to 'two main-line items' and the pricing-page source removed, with a note at the top; the same list in the English edition and 'four items' in the Japanese edition were corrected the same day with their own notes. The email sent on 9/13 is the original.
Editions affected: 2026-09-13
-
Corrected 2026-09-15Japanese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- The breakdown sentence in the Japanese edition's 'Sources & Inventory' section said the body used 16 receipts, of which 3 came from the previous night's intake and the other 15 were fetched separately; 3 plus 15 is 18, which does not match the 16 in the same sentence.
- What was actually the case
- The body actually cited 16 external receipts: 3 from the previous night's intake and 13 fetched separately or drawn from what we already hold.
- Which layer failed
- Bad data — the publish-time recount only syncs the phrasings registered in its position table; the inventory section's follow-on, anaphoric and arithmetic-breakdown sentences are not in that table, so only the opening line and the footer got synced and the rest stayed at their generation-time values.
- What we did about it
- The 15 in that sentence was corrected to 13 on the Japanese archive page, the list of sources that follows was left untouched, and a note was added at the top of the page; the same error in the Chinese edition is logged separately.
Editions affected: 2026-09-14
-
Corrected 2026-09-15Japanese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- The Japanese edition's 'Sources & Inventory' section put the count of external sources cited in the body at 17, while that edition's footer and opening line both read 15.
- What was actually the case
- The body actually cited 15 external sources; the 17 was a stale generation-time value.
- Which layer failed
- Bad data — the publish-time recount only syncs the phrasings registered in its position table; the inventory section's follow-on, anaphoric and arithmetic-breakdown sentences are not in that table, so only the opening line and the footer got synced and the rest stayed at their generation-time values.
- What we did about it
- The Japanese archive page's 17 was corrected to 15 and a note added at the top of the page; the same error in the Chinese edition is logged separately.
Editions affected: 2026-09-04
-
Corrected 2026-09-15Japanese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- The Japanese edition's 'Sources & Inventory' section put the count of external sources cited in the body at 28, while that edition's footer and opening line both read 26.
- What was actually the case
- The body actually cited 26 external sources; the 28 was a stale generation-time value.
- Which layer failed
- Bad data — the publish-time recount only syncs the phrasings registered in its position table; the inventory section's follow-on, anaphoric and arithmetic-breakdown sentences are not in that table, so only the opening line and the footer got synced and the rest stayed at their generation-time values.
- What we did about it
- The Japanese archive page's 28 was corrected to 26 and a note added at the top of the page; the same error in the Chinese edition is logged separately.
Editions affected: 2026-09-03
-
Corrected 2026-09-15Chinese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- This edition's 'Sources & Inventory' section put the count of external sources cited in the body at 17, while the footer and the opening line both read 15.
- What was actually the case
- The body actually cited 15 external sources; the 17 in the inventory section was a stale generation-time value.
- Which layer failed
- Bad data — the publish-time recount only syncs the phrasings registered in its position table; the inventory section's follow-on, anaphoric and arithmetic-breakdown sentences are not in that table, so only the opening line and the footer got synced and the rest stayed at their generation-time values.
- What we did about it
- The archive page's 17 was corrected to 15 and a note added at the top of the page; the same error in the Japanese edition is logged separately.
Editions affected: 2026-09-04
-
Corrected 2026-09-15Chinese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- This edition's 'Sources & Inventory' section put the count of external sources cited in the body at 28 and said it was the same figure printed in the footer, while the footer and the opening line both read 26.
- What was actually the case
- The body actually cited 26 external sources; the 28 in the inventory section was a stale generation-time value.
- Which layer failed
- Bad data — the publish-time recount only syncs the phrasings registered in its position table; the inventory section's follow-on, anaphoric and arithmetic-breakdown sentences are not in that table, so only the opening line and the footer got synced and the rest stayed at their generation-time values.
- What we did about it
- The archive page's 28 was corrected to 26 and a note added at the top of the page; the same error in the Japanese edition is logged separately.
Editions affected: 2026-09-03
-
Corrected 2026-09-15Chinese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- The 9/14 zh edition's inventory sentence carrying an arithmetic breakdown put the external-receipt count at 18, split as 3 from the previous night's intake and 15 retrieved separately, disagreeing with the 16 printed in the same issue's opening line and footer.
- What was actually the case
- The body actually cited 16 external receipts: 3 from that night's intake and 13 retrieved separately or pulled from the archive. The English and Japanese editions of the same day both printed 16.
- Which layer failed
- Bad data — the publish-time recount syncs only the phrasings registered in its position table; an arithmetic breakdown inside the same sentence ("only N ... the other M ...") is not in that table and stayed at its generation-time value.
- What we did about it
- On 2026-09-15 the two figures in that sentence on the archived page were corrected to 16 and 13, with a correction note added at the top of the page; the list of enumerated sources that follows was left untouched. The email sent on 9/14 is the original version and cannot be recalled; nothing else in the issue was changed.
Editions affected: 2026-09-14
-
Corrected 2026-09-15Chinese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- The 9/11 zh edition's "Sources & Inventory" section gave the external-receipt count as 14 in three places, and two back-referencing sentences repeated 14, while the same issue's opening line and footer printed 15.
- What was actually the case
- The body actually cited 15 external receipts (counted as deduplicated public-version links minus our own domain). The three inventory-section figures and the two back-references were stale generation-time values; the opening line and footer were right. The English and Japanese editions of the same day both printed 15.
- Which layer failed
- Bad data — the publish-time recount recognized only the opening-line and footer phrasings at the time; the three inventory-section phrasings and back-references of the form "those N receipts" were not in its position table, so only two spots were synced and the rest stayed stale.
- What we did about it
- On 2026-09-15 the five figures on the archived page were corrected to 15 and a correction note was added at the top of the page. The three inventory phrasings were added to the recount's position table on 2026-09-14; the back-reference gap is tracked on an internal card. The email sent on 9/11 is the original version and cannot be recalled; nothing else in the issue was changed.
Editions affected: 2026-09-11
-
Corrected 2026-08-13English editionBad inference — the inputs were right, the reasoning or the wording was not
- What we wrote
- The 8/10 English edition's opening sentence folded July's break-in at Hugging Face together with the internal message board disclosed at Black Hat, which made the security expert we quote further down read as if he were denying the break-in.
- What was actually the case
- The break-in itself was reported correctly — a zero-day plus stolen credentials is not access anyone granted (jointly disclosed 2026-07-21). What was wrong was fusing two separate incidents into one sentence, manufacturing a contradiction that was never in the reporting.
- Which layer failed
- Bad inference — two events compressed into a single sentence, and the sentence structure itself generated a conflict the sources did not contain.
- What we did about it
- On 2026-08-13 a dated correction note was added at the top of the English archive page (what it said, what changed, why); the opening sentence was rewritten to separate the two incidents, and a contradiction marker was appended to the security-community line. The email already sent is not recalled, and the note says so. We also judged this failure type machine-catchable in a narrow form, and routed the rule to a follow-up ticket.
Editions affected: 2026-08-10|Read that edition
-
Corrected 2026-08-13Chinese editionBad inference — the inputs were right, the reasoning or the wording was not
- What we wrote
- The 8/10 zh edition called Hugging Face "the injured party" without ever saying what injury it suffered; in the same passage a security-community line ("they never had to escape the sandbox") sits head-on against OpenAI's own account of the July incident, and we printed both without saying they collide.
- What was actually the case
- The injury refers to a separate July incident: OpenAI's model used a zero-day in a package-registry cache proxy to reach the open internet, then used stolen credentials to break into Hugging Face's production database and take benchmark answers; Hugging Face detected and contained it before OpenAI told them (jointly disclosed, 2026-07-21). The two accounts of "escape" really do conflict, and our own rule says conflicts get flagged.
- Which layer failed
- Bad inference — a background event was compressed into an unexplained label, and we broke our own rule that keeping a contradiction means marking it, not printing each side once.
- What we did about it
- On 2026-08-13 the website archive page was amended in place with a correction-and-addendum note, and the contradiction marker was added to that passage. The note states plainly that the email sent on 8/10 was the original version and cannot be recalled, and that nothing else in the edition was touched.
Editions affected: 2026-08-10
-
Corrected 2026-08-11Chinese editionBad inference — the inputs were right, the reasoning or the wording was not
- What we wrote
- Our 8/9 and 8/10 editions both wrote that the company "wiped and rebuilt" the repository, which reads as if OpenAI found the agent message board first and then decided to wipe it.
- What was actually the case
- Per a timeline correction posted on the evening of 2026-08-09 by an engineer self-identifying as being on OpenAI's side: when the repository hole was first patched the company did not know the message board existed; the board was cleared incidentally during the rebuild. The first action actually aimed at the board came later.
- Which layer failed
- Bad inference — our sentence collapsed "patched the hole" and "acted on the board" into one move, implying a decision sequence we had no evidence for.
- What we did about it
- The 2026-08-11 edition ran the correction as a main item and, in the same breath, listed three weaknesses in the correction itself (single source, posted from a personal account, and self-serving in direction). The two outside commentators who read the evidence the other way were left standing; we did not adjudicate between them.
Editions affected: 2026-08-09, 2026-08-10
-
Corrected 2026-07-11Chinese editionBad inference — the inputs were right, the reasoning or the wording was not
- What we wrote
- An earlier edition presented two papers together as "the two seed papers of the reasoning narrative".
- What was actually the case
- The second one is architecture-evaluation methodology and has nothing to do with reasoning. They should never have been paired.
- Which layer failed
- Bad inference — the two arrived in the same batch on adjacent topics and got grouped without each being checked back against its own abstract.
- What we did about it
- The 2026-07-11 edition split them in an explicit corrections paragraph, giving each paper its own takeaway.
Editions affected: 2026-07-09
-
Corrected 2026-07-09Chinese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- Our coverage of the three federal pressure moves against Anthropic never said how they ended in court.
- What was actually the case
- All three had been blocked by the courts back in March 2026. That judicial brake was there the whole time; we only caught it that day.
- Which layer failed
- Bad input by omission — our retrieval covered executive-branch sources but not the court rulings from the same period.
- What we did about it
- The 2026-07-09 edition carried the fix in the headline rather than burying it, and court rulings became a mandatory source for this topic.
Editions affected: 2026-07-08
-
Corrected 2026-07-08Chinese editionBad input — what came in was already wrong, and our checks missed it
- What we wrote
- Two earlier editions described the U.S. government's Fable ban as still in force, and reasoned about government leverage on that basis.
- What was actually the case
- The ban had already been lifted on 2026-06-30 — it ran 18 days and was replaced by standing pre-release review plus a dedicated team.
- Which layer failed
- Bad input. We had no rule requiring us to chase a contradiction to the end when reporting collides with observable reality — and the report in question was itself generated with Fable 5.
- What we did about it
- The whole 2026-07-08 edition ran as the correction: what was wrong, the root cause, and the three rules we added because of it. The same edition struck one entry from the prediction ledger: Lambert's "government and Anthropic will settle" was not an ex-ante call, since the lifting was already under way when he wrote on 7/1.
Editions affected: 2026-07-06, 2026-07-07