Daily Brief SecondSource Morning Brief · October 5, 2026 · Oct 5, 2026
Why today matters: of the eight preconsidered bills the City Council hears today, the heaviest would require a model to pass third-party validation before it goes on the market, and require that a person be able to shut it down. The day before, Sam Altman said in an interview that the world should accept some bad things happening in exchange for AI's benefits; the interview did not mention the bill.
1. New York City Council hears a bill not yet introduced: AI models without third-party validation, or that a human can't shut down, couldn't be sold in the city. (Affects: companies selling AI in New York)
2. OpenAI CEO Sam Altman says his stance on regulation differs a lot from Anthropic's; OpenAI's former head of policy research says the two labs' practices barely differ. (Affects: anyone comparing the two labs, who shouldn't rely only on what the CEO says)
3. The US leads 17 countries, itself included, in endorsing a Kyoto science declaration; the White House's description uses its new name for AI, "super intelligence," and the press release lists no funding or binding terms. (Affects: researchers in the endorsing countries, though nothing is funded or binding yet)
None of the claims we screened this week was strong enough to go on our tracking list. This issue instead goes back to three of our own judgments that already have answers and that the column hasn't written up yet. We checked and closed all three on September 30. The Sonnet 5 judgment's verdict date was September 1. The Microsoft annual-report judgment and the August 16 judgment, which asked about prices on September 1, both came due on September 15.
The short version: Sonnet 5's US$2 input and US$10 output is now Anthropic's standard price, so two of our judgments were wrong, and a third, on Microsoft's annual report, can't be settled from the source we named.
The receipts: would Anthropic's Sonnet 5 go back to full price on September 1?
① What we said then: we opened the entry on July 31, meaning we started tracking a judgment still to be verified. The original wording: "After the Anthropic Claude Sonnet 5 API promotional price expires on 2026-08-31, the list price on the official pricing page returns to the standard pre-promotion list price," that is, US$3 per million input tokens and US$15 per million output tokens. A token is the unit AI models count text in — roughly a few characters each — and usage is billed per token. The introductory price at the time was US$2 input and US$10 output, so returning to standard meant a 50% increase. Our confidence was 0.85 (out of 1).
② The actual reading: the source we named in advance was Anthropic's official pricing page. Its note 3 says Sonnet 5's price of US$2 input and US$10 output, "announced at launch as introductory pricing through August 31, 2026, is now the standard price," and the increase originally scheduled for September 1 "will not occur." We checked on September 30 and again on October 4; the text was the same (Anthropic pricing page).
③ Ruling, and where it went wrong: for teams already costing out Sonnet 5, US$2 input and US$10 output is now the official standard price, not a promotion that expires. But this entry also shows that a standard price can be changed by the vendor too. We got it wrong because events overturned our premise: the increase had been scheduled, and Anthropic then made the promotional price the standard price. The pricing page now says the increase won't happen; we did not find exactly when the change was made. Our 0.85 rested entirely on "scheduled terms will run as written" and left no room for the company changing its mind.
④ Which of our tracked judgments moved: none yet. It feeds two judgments we track: whether frontier AI companies can hold their list prices, and whether AI companies can keep their margins once they sell usage. Both keep their old confidence for now and haven't been adjusted for this reading.
⑤ Does our standard change: yes, in one respect. Because a company can reprice on its own before a deadline, from now on when we judge that "a scheduled price change will go through," we cap our own confidence at 0.7, and we list "a repricing announced before the deadline" in advance as a case that counts against us.
The receipts: would Microsoft's annual report say whether OpenAI's cloud credits were counted as Azure revenue?
① What we said then: the column opened this entry on August 2. Microsoft invested in OpenAI partly in cloud credits, and OpenAI then used those credits to buy Microsoft's Azure cloud services. If the credits are counted as Azure revenue, the money Microsoft invested comes back onto its books as customer revenue. Bill Gurley of the venture firm Benchmark said the credits were counted as Azure revenue; Microsoft CEO Satya Nadella said they were not. "One of them has to be wrong," and the accounting notes in Microsoft's fiscal 2026 annual report could settle it. We recorded no confidence at the time.
② The actual reading: the source we named in advance was Microsoft's fiscal 2026 annual report (10-K, filed with the US Securities and Exchange Commission on July 29, 2026, accession number 0001193125-26-323660). Reading it directly on September 30, the notes say "we recorded revenue from commercial arrangements with OpenAI, inclusive of revenue-sharing payments," amounting to US$24.1B. The notes give only the total; they don't say how the credits were booked.
③ Ruling, and where it went wrong: unresolved. The source we named in advance can't answer the question, and before opening the entry we hadn't checked whether past annual reports ever disclosed this kind of figure.
④ Which of our tracked judgments moved: none. Gurley's and Nadella's claims stay at their original confidence; neither has been confirmed or overturned. For readers, this means what Microsoft made public in its fiscal 2026 annual report can't settle the dispute over whether the credits count as Azure revenue.
⑤ Does our standard change: yes, in one respect. Before naming a source for an entry, we first confirm that the source has disclosed the same kind of figure before; if we can't find a precedent, we name a backup source as well.
The receipts: was Anthropic "afraid to raise prices"? Did the price go back to standard on September 1 as the terms said?
① What we said then: our August 16 column judged that Gurley's claim — "nobody dares raise prices; Anthropic won't touch prices for fear of losing share" — did not hold. The hardest evidence was the official pricing page stating that Sonnet 5's introductory price would expire on August 31. So we tracked whether the price returned to standard on September 1 as the terms said (a 50% increase). We recorded no confidence at the time.
② The actual reading: the same pricing page and the same note 3 as the first entry: US$2 input and US$10 output is now the standard price, and the increase will not occur.
③ Ruling, and where it went wrong: we got it wrong, and this time the fault was our reasoning: we treated an expiry date as evidence that the company dared to raise prices, but a company can change terms even after writing an expiry date into them. The strongest plank of our August 16 rebuttal of Gurley is therefore void. The receipts on that claim have been changed from "doesn't hold" back to "unresolved," and notes have been added to that column and to that day's morning brief.
④ Which of our tracked judgments moved: on October 3 the column's ruling changed from "doesn't hold" to "unresolved." Three tracked judgments depend on it: what our recorded facts on AI pricing and market share now say, whether a price war will happen, and how far software-industry profits can be squeezed at most. We haven't re-evaluated any of the three.
⑤ Does our standard change: not separately. The mistake is the same as in the first entry — reading scheduled terms as a statement of the company's resolve — and the first entry already changed the standard.
The receipts page for this entry is in Chinese only for now.
All our judgments and verdict dates → the judgment index
Why this matters to you: if your model service has users in New York, you can check two things now: whether a person can shut it down, and whether it can be handed to an outside validator. The bill's scope isn't defined yet, and nothing here is law today.
Legistar, the City Council's official legislative record system, lists a full Council meeting at 11 a.m. today. The first item is an oversight hearing, "Examining the Risks Posed by Artificial Intelligence," followed by eight preconsidered AI bills. A preconsidered bill is one that hasn't been formally introduced yet and is brought to a hearing for discussion first. The bill pages list a formal introduction date of October 8, and there is no vote today, so none of these is law (New York City Council Legistar, 2026-10-05 agenda). The official summary of the heaviest bill reads: "This bill would make it unlawful to market, offer for sale, sell, or deploy an artificial intelligence (AI) model in New York City that has not received third-party validation or does not have a technical capability to be shut down by a human operator." The validator would have to deliver its conclusion — whether the model passed validation and whether it can be deployed — to the developer and to Cyber Command, the city's cybersecurity agency. Validators would also have to disclose their own interests in the model, and each violation carries a US$25,000 penalty. Another bill would let anyone report AI violations to the city's Department of Consumer and Worker Protection; the reporter could receive 25% of the amount recovered, or 50% if they serve or bring the case themselves.
In a September 16 press release, the Council said Speaker Julie Menin had invited Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman to appear in person (New York City Council, 2026-09-16). The local outlet amNewYork reported on September 28 that OpenAI, Google, Anthropic and Meta agreed to send representatives to testify under oath, not their CEOs; Elon Musk's AI company didn't respond to the Council and was subpoenaed by the Speaker (amNewYork, 2026-09-28). Status as of this issue: the hearing hasn't started, and we don't yet know who appears or what they say.
Verification: for the agenda and the bills, we read the summaries on the Council's official pages, not the full bill text; we pulled those summaries with an automated tool and no person checked them word for word. That the four companies agreed to testify under oath and that Musk's company was subpoenaed has so far been reported only by amNewYork. ⚠️ The summary doesn't define what counts as an "AI model": only the most advanced large models, or everything? And if users in New York City call a model hosted elsewhere through the cloud, does that count as "deploying" it in the city? Those two points decide whether the bill binds AI labs at all. The summary also doesn't say whether "shut down" means stopping the whole service or halting a single run, or how it would apply to open-weight models, which anyone can download and run themselves. ⚠️ Gary Marcus, a commentator who has long argued for regulating AI, previewed testimony for the hearing and called the bill that pays reporters a share of recoveries a whistleblower bill (Gary Marcus, 2026-10-05); on the official agenda, whistleblower protection is a separate bill that covers city contractors, and we go by the official agenda.
Judgment update: our October 4 deep dive (published in Chinese and Japanese only) looked at the White House AI accord of September 29. That accord is a voluntary document signed at the White House by the heads of Google, OpenAI, Anthropic, Meta, xAI, NVIDIA and others, with no legal force; the deep dive argued that what it really leaves unwritten is two things: who draws the audit's scope, and who gets to see the results. Going by the official summary, the New York bill answers the second: results go to the developer and the city, and validators must disclose their interests. Who draws the scope, the summary doesn't say. We haven't read the full bill text. Our October 3 issue logged six accountability moves, by Senator Hawley, state authorities in Florida, California and New Mexico, the FTC and one private lawsuit. All rely on existing law and, except for one Florida motion still awaiting a court ruling, target conduct that has already happened; none sets rules a model must meet before going to market. Separately, according to The Wall Street Journal as relayed by the financial site Investing.com, the first article of the charter of the AI task force the White House set up on October 3 says it will not adopt "regulation that could restrict innovation or competition" (Investing.com, 2026-10-03). With the federal side signaling up front that it won't set rules of this kind, the first written rule on who must sign off before a model reaches the market comes from a city. So we are logging a new working read: cities or states, not the federal government, may be first to write that sign-off into law. The evidence is weak: one city, and one bill not yet voted on.
What would prove this wrong: if this validation bill doesn't get out of committee by the end of this year, and no other state or large city proposes the same provisions, this working read weakens sharply. If the federal government expressly declares this year that localities may not set their own rules on putting AI on the market, we rewrite it as "the federal government took back the decision over whose sign-off a model needs." The nearest thing to watch is the version formally introduced on October 8: whether it adds a definition of "AI model," and whether it counts cloud calls. The verdict date is one we set ourselves: December 31, 2026.
Why this matters to you: when you compare the two labs, don't rely only on what the CEO says. Altman says they differ a lot on regulation; Brundage says four product and deployment practices are the same at both, a claim we haven't independently checked. The two may not be talking about the same thing.
On October 4 the tech outlet SiliconANGLE cited an interview in Politico's new newsletter, Decoded, in which Altman said "I think there's a lot of daylight." He said he doesn't accept a standard under which "we'll make sure there's no major hacks, there's no misuse of this technology, there's zero scams," and he called concentrating powerful AI in a single lab a "completely unacceptable trade-off" (SiliconANGLE, 2026-10-04). The sentence Politico reporter Jonathan Martin quoted: "we believe that the world should accept some bad things happening for the benefits of this technology and people having the agency" (Jonathan Martin, 2026-10-04).
The same day, Miles Brundage, formerly head of policy research at OpenAI and now independent, pushed back: "There really isn't a lot of daylight." His reasons: both labs build some values into their products, both rarely publish model weights, both have huge numbers of free users, and both give their best security tools to defenders first. He closed with "Narcissism of small differences" (Miles Brundage, 2026-10-04). He also said the Anthropic example Altman gave when pressed didn't match the facts.
Verification: we didn't get the Politico original. We read SiliconANGLE's page; Forbes and Reuters reported on the same interview that day, and we read only their headlines. All of them trace back to the same interview, and what can be confirmed is that Altman said these things. We read Brundage's post in the original. ⚠️ We have not independently checked Brundage's claim that Altman's example was inaccurate. ⚠️ Brundage used to work at OpenAI and has positions on both labs; our analysis was produced with help from Anthropic's models, and Anthropic is one of the parties in this item.
Judgment update: we keep both claims on the books. Our reading, not either man's words: Altman seems to be talking about public positions on regulation, while Brundage lists product and deployment practices, so the two may not be talking about the same thing. At today's hearing, watch whether the two labs' representatives take different positions on the bill's third-party validation and human-shutdown requirements: a split would be evidence, at the level of practice, for the gap Altman describes; matching positions would fit Brundage's reading. Read together with main-line item 1: according to amNewYork, OpenAI agreed to send a representative to testify under oath at today's hearing; Altman's interview didn't mention the New York bill.
Why this matters to you: for researchers in the endorsing countries, nothing changes yet: the press-release summary lists no funding figures or binding terms. What is new is wording. The White House used its new term, "super intelligence," to describe a declaration endorsed by 17 countries, the US among them; we haven't read the full text, so we can't say whether the declaration itself uses the term.
Michael Kratsios, director of the White House Office of Science and Technology Policy, announced on October 4 that at the Science and Technology in Society (STS) Forum, an annual science-policy conference held in Kyoto, the US led 17 countries, itself included, in endorsing the "Kyoto Vision for a Golden Age of Science": "The way we make scientific discoveries has changed, and the Kyoto Vision is a recognition of that shift and a call to action." (Michael Kratsios, 2026-10-04) The White House press release sets out three components: reforming how research institutions are funded and organized; integrating super intelligence into the research process and widening researchers' access to the relevant tools, scientific data, compute and experimental facilities; and investing in students and early-career researchers selected on merit. The endorsing countries are Argentina, Bulgaria, Chile, Cyprus, Germany, Greece, Indonesia, Italy, Japan, Kazakhstan, South Korea, New Zealand, Poland, Singapore, the UAE, the United Kingdom and the United States (White House, 2026-10-04). "Super intelligence" (SI) is the new name a September 29 White House executive order gave to what government language had called "AI" (White House, 2026-09-29).
Verification: the announcement has two first-hand sources, Kratsios's own post and the press release on the White House site, but both come from the same publisher; we read the press release's summary, not its full text word for word. ⚠️ France and Canada are not on the list, nor are China or India; the sources don't say whether they weren't invited, declined to sign, or are on a different timeline.
Judgment update: no judgment changes; we log one reading. Within what we track, this is the first time since the renaming order that the US has used the term "super intelligence" in describing a document endorsed by several countries; the wording comes from the White House press release. What we read contains only a framework and terminology; we saw no funding figures or binding terms. It would gain substance only if an endorsing country later attaches a budget or compute program to the declaration.
What to take away today: #1: if your model service has users in New York, you can check two things now: whether a person can shut it down, and whether it can be handed to an outside validator; the bill's scope isn't defined yet, and nothing here is law today. The other items: nothing to act on today.
1. [Today] (relayed October 4) According to documents seen by Bloomberg, a financing company controlled by a Chinese local government funded a listed company's purchase of more than 700 servers, 32 of them ASUS machines carrying NVIDIA's B300 data-center GPU; under US rules, the B300 can't be sold to China without a license. We haven't read the Bloomberg original (Poe Zhao, 2026-10-04; Teortaxes, 2026-10-04).
2. [This week] (reported October 3) According to Axios, Reflection, a US AI company backed by NVIDIA, is preparing to release an open-weight model (one anyone can download and run themselves); sources expect it to compete with the top Chinese open models while initially trailing the strongest US systems. The anonymous commentator account Teortaxes, which reposted it, is skeptical. We haven't read the Axios original (@choblin29, 2026-10-03; Teortaxes, 2026-10-04).
1. [This quarter] (specs disclosed August 26, interview relayed September 30, picked up by us today) Specs for OpenAI's in-house inference chip, Jalapeño: six latest-generation HBM4 memory stacks per package, 216 GiB, 15.4 TB/s of bandwidth, TSMC 3nm, rated peak 700 W. To size the HBM demand from OpenAI's in-house chip, start with six HBM4 stacks per chip, then multiply by shipment volume; no production or shipment figures are public. An inference chip handles the half of the work that runs a model after deployment, answering queries. Our August 31 issue covered its output per unit of power and total cost of ownership against NVIDIA racks, and our October 4 issue covered how OpenAI's head of hardware, Richard Ho, used models to design it. The 700 W and the total cost of ownership were covered before; what's new today is the memory configuration and system scale. TrendForce, a Taiwanese industry research firm, writes that each package places the compute die alongside six HBM4 stacks, "delivering 216 GiB of memory and 15.4 TB/s of bandwidth," on TSMC's 3nm process (TrendForce, 2026-08-26). HBM is high-bandwidth memory: several layers of memory stacked vertically and packaged together with the compute die. Semiconductor analyst Ian Cutress's newsletter More Than Moore adds a rated peak of 700 W and a measured sustained draw of close to 550 W, with measurement conditions undisclosed; 128 accelerators form one interconnected unit (a "local domain"), and the full system has 2,048 (More Than Moore, 2026-09-30). ⚠️ Both pieces relay OpenAI's own disclosure, and we didn't read OpenAI's original page, so this is two channels carrying one claim, with no third-party measurement; the 550 W is OpenAI's own measurement; the system scale appears only in Cutress's piece. TrendForce also says the HBM4 "is believed to" be supplied by Samsung; we don't use that. ⇒ Our working read, not settled: to whatever extent an in-house chip substitutes for NVIDIA, it substitutes for compute silicon, not memory; each package still needs six HBM4 stacks bought from memory makers, so its effect on HBM demand scales with shipment volume, which is undisclosed.
1. [Today] (posted October 4) Vercel CEO Guillermo Rauch: once AI agents take over writing code, the return on switching programming languages has to be recalculated. Vercel is a website deployment and front-end cloud platform. Rauch looked back at Vercel's early rewrite of its build tool Turborepo from Go into Rust; Rust is known for guaranteeing memory safety at compile time, but it takes more effort to write. The rewrite was finished technically, but with people writing it line by line it was expensive, and there was heavy internal debate over whether it was worth it. He argues that once agents write the code, the cost structure changes: the language that's easiest for people to write is no longer the same as the language that's best for the business (Guillermo Rauch, 2026-10-04). ⚠️ There are no figures for the cost or benefit of the rewrite; Vercel sells agent infrastructure, so "agents change everything" fits its business; he doesn't say whether the cost of human review and testing goes down after an agent does the rewrite. ⇒ One thing you can measure yourself: next time you evaluate a rewrite, estimate the hours for "agent rewrites, people review" separately from "people rewrite it themselves."
1. [This week] (paper online October 1) A study turns "pick which older paper can inspire new research" into a test: letting AI search on its own over multiple steps did worse than plain vector-similarity retrieval. The research team asked 184 first authors of computer-science projects to mark which prior papers genuinely advanced, or could have advanced, projects they had already completed; the systems could search only the literature available when each project began. Measured by how often the right paper appeared in the top 20, agentic search, where the model searches, reads and then decides its next step on its own, scored 0.42; retrieval that turns text into vectors and compares similarity scored 0.48; the best system scored 0.51, and the abstract doesn't say which kind it is (ScholarCatalyst, arXiv 2610.02202). A score of 0.48 means that on roughly half the questions, the right paper landed in the top 20. Co-author Chelsea Finn, a Stanford professor, flagged the paper on X on October 4 (Chelsea Finn, 2026-10-04). ⚠️ We read only the abstract; what's measured is the authors' after-the-fact judgment, not how inspiration actually happened; the sample is computer-science projects only; the abstract doesn't say which model ran the agentic search or how it was configured. ⇒ On this one computer-science benchmark, letting a model search over multiple steps on its own lost to plain vector retrieval; teams building their own literature-search tools can start with vector retrieval as the baseline to beat. The agentic loss may come down to tooling or model configuration, which the abstract doesn't cover, so it can't be extended to agentic methods as a whole.
2. [This week] (paper online September 29) A preprint has the model write itself a one-line lesson after each wrong answer and retry with that lesson in hand, and it learns on problems where standard reinforcement learning learns nothing at all. Reinforcement learning here means having the model answer repeatedly and rewarding or penalizing it based on automatically verified results. The common GRPO method gets its signal by comparing several attempts at the same problem; if every attempt is wrong, there's no signal to learn from. The authors picked tool-calling and coding problems on which Alibaba's open model Qwen 3.5 9B Thinking got all 128 attempts wrong. After standard GRPO training, its first-try accuracy stayed at 0 to 1%. The new method, RLTL;DR, has the model read the verification result and write a one-line lesson, then answer with that lesson during training; first-try accuracy rose to 14% to 31%, and with the lessons removed at evaluation it still reached 12% to 13% (RLTL;DR, arXiv 2609.37633). ⚠️ A single preprint, tested on one 9-billion-parameter open model only, on problems the authors chose, applicable only to tasks whose answers can be checked automatically, with no independent replication. ⇒ For teams post-training their own models on tasks with automatically checkable answers, this offers a route to try even on problems where every attempt fails. But the lesson amounts to giving the model an extra line of text feedback during training, so the information available differs from standard reinforcement learning; it can't yet be used to rebut the claim that "reinforcement learning only draws out what a model already knows," only to add an example run under different conditions.
No product news this issue. Every product-company article we read had an original publication date of September 16 or earlier, older posts resurfacing; another 170 pieces went unread, most of them older posts backfilled from the official blogs of xAI, Cohere and Perplexity, and we haven't checked one by one whether any new product article is among them.
1. [Look back] (deep dive, September 17, 2026) To compare AI model prices, switch from list price to "cost per task" — but that ruler itself moves. Within one week, the same independent evaluator, Artificial Analysis, revised its numbers so that Anthropic's flagship Fable 5.1 went from "20% more expensive per task than its predecessor" to "13% cheaper," without its list price changing at all; the difference was that the scale moved from the September 1 version to v4.3, and "20% more expensive" held only under the September 1 version. Where the two vendors really diverge is in who pockets the tokens saved (the unit AI usage is counted and billed in). OpenAI's GPT-6 Astra, at the same composite score, uses about 27,000 output tokens per task against about 78,000 for Fable 5.1, but Astra's list price is set at 2.5 times that of its own predecessor, GPT-5.6 Sol; the deep dive judged that part of the advantage of using fewer tokens is clawed back through list price. The token counts are compared with Fable 5.1 and the list price with OpenAI's own predecessor, so the bases differ, and cost per task can't be computed from the piece.
What you can do: take last month's bill and look at the ratio of cache reads to output; cache reads are the lower rate charged when a repeated prompt prefix hits the cache. The September 17 deep dive's judgment: cost per task in an evaluation measures standard work, while your bill measures your own usage; the higher your share of cache reads, the more changes in cache pricing affect you, so that ratio may say more than evaluation numbers about whether you'll pay more or less after a generation change. The full deep dive is published in Chinese and Japanese only; there is no English edition.
Follow-up: on September 22 Anthropic released Opus 5.5, saying it reaches Fable 5.1's level on most work and that running typical work costs 40% less than on Opus 5; Anthropic didn't explain how the 40% was calculated, and it is the company's own measurement, not a third party's (Anthropic, 2026-09-22).
Each day we hunt the AI firehose for the insights that matter and the practitioner judgments worth tracking over time, show how every item was verified, and ask which judgment got harder and who's been right.
— SecondSource · generated by our research system · 23 sources · Got a view? Reply and tell us
Written from the same research and judgments as the Traditional Chinese edition. Sources are linked; we distinguish original documents from reporting and mark what we could not verify.