Daily Brief SecondSource Morning Brief · September 18, 2026 · Sep 18, 2026
1. None of the three sets of numbers circulating under "pacing the frontier" is a speed. OpenAI cut its flagship line's compute by nearly 60%, but the total fell only about 2%. (Affects: investors sizing AI compute demand)
2. CoreWeave says its newly signed short contracts price at roughly five times its own fleet average, while public GPU rental prices rose just 5% in a year. (Affects: CFOs buying compute)
3. Put the same model in a different harness and the cost per solved task can differ by up to 5.1x. The harness is the layer of software wrapped around a model that manages its files and tools. (Affects: engineering leads rolling out AI coding tools)
This issue draws on the research report written early on September 18, 2026, and a deep-dive column published the same day; the material spans June 27, 2026 to September 18, 2026. Last night's sweep covered 236 pieces, and 17 clickable outside receipts made it into this issue. This is the email edition; the full edition of this issue is the archive of record.
Why this matters to you: the next time you read that an AI lab "slowed down" by some amount, ask what the denominator is and who measured it before deciding whether to believe it.
Core judgment: our September 15 issue said nobody had attached a speed number to "pacing the frontier," the idea floated by Anthropic chief executive Dario Amodei, and that the only thing outsiders could check was whether evaluators got in the door. Today's deep-dive column rechecked that and narrowed it: there are three sets of numbers, and none of them says how fast capabilities may advance per year. The first comes from AI Futures Project, a research group that forecasts AI progress. On August 5 it proposed that frontier companies spend at least about 70% of their compute serving outside customers and at least 25% on safety research whose results are fully published, with the shares "enforced by third-party auditors with extensive internal access in AI companies" (AI Futures Project, 2026-08-05). The second is the reinforcement-learning compute allocation OpenAI published for itself on September 6 (OpenAI, 2026-09-06); reinforcement learning is the main stage after pretraining where a model's capabilities get pushed up. The third is the "30-day safety stand-down" called for by George Whitesides, a US congressman from California. We read it only as a repost by AI policy commentator Zvi Mowshowitz (Zvi Mowshowitz, 2026-09).
We covered the three percentages in the second set in our September 7 issue: in the week after August 7, GPU allocation for the Astra-class flagship headed for release fell 59.2%, other model categories rose 17.2%, and that offset about 85% of the cut. What's new today is the conversion. Work backward from those three numbers. If a 17.2% rise in the other lines made up 85% of a 59.2% cut in the flagship, the other lines must have had about three times the flagship's allocation, so the flagship started with only about a quarter of the reinforcement-learning compute analyzed. Spread the unrecovered part of the cut across the whole, and total compute after the brakes went on fell only about 2%. That's our own calculation. It assumes the 85% is a ratio of allocation volumes and that both percentages share a base period; the original wording supports that reading but doesn't spell it out. The same event can be written as "flagship training compute cut by nearly 60%" or as "total reinforcement-learning compute barely moved." Only the denominator differs.
Why we dug in now: the September 15 judgment split the pledge into two halves, access that can be checked and speed that can't. The column recasts them as one thing. The proposal's three levers are training compute, the kind of training, and using AI to improve AI. From outside, you can at most see total compute, not how it's split by use. So as long as pacing is pinned to the development side, someone has to go inside to read the speed. Why doesn't anyone write down a speed? The column found a more concrete possible reason. Antitrust adviser Doug Calidas told The Washington Post's technology policy brief on September 15: "If they agree to restrict output, that's a pretty classic antitrust violation" (The Washington Post, reposted on Senator Jim Banks's official site, 2026-09-15). The goodwill that Stanley Woodward, the US deputy attorney general, signaled on September 17 covers only cybersecurity cooperation (Washington Examiner, 2026-09-17). Federal Trade Commission chair Andrew Ferguson, meanwhile, said that companies asking for regulation and for an antitrust exemption at the same time set off his alarm bells (Reuters, as relayed by silicon.co.uk, 2026-09).
Open ?
Verification: we read the AI Futures and OpenAI originals directly. We saw Whitesides only in a repost, and we did not read Woodward's remarks or the original Ferguson report. ⚠️ OpenAI's allocation is self-reported with no third-party sign-off, and it covers only the reinforcement-learning slice over those few weeks, not all of the company's compute. ⚠️ "Writing down a speed number is legally closest to competitors agreeing to restrict output, so nobody does it" is our own inference, and a soft one: Calidas was describing legal risk, and no lab has said that's why it doesn't publish numbers. ⚠️ Disclosure: our research system runs on Anthropic's models; on the Anthropic parts of this item we only relay sources and structural analysis.
Judgment update: the column concludes that access isn't a substitute for speed; it's the only ruler for measuring development-side speed, and nobody has yet said which number it should read. The ruler itself is still just a promise. TechCrunch reported on September 16 that neither Anthropic nor OpenAI has said which evaluators they'll work with, when they'll start, how many there will be, or what they'll see. Adam Gleave of FAR.AI, a nonprofit that does AI safety evaluations, said his group has turned down engagements because a frontier developer wanted too much control over the evaluation process (TechCrunch, 2026-09-16). We haven't decided whether to formally rewrite our long-running judgment along the column's lines. If you read only this far: whenever you see "lab X slowed down by Y," ask three things. Is the denominator one model line or the whole company? Who measured it? Has the person measuring ever been inside? Today's answers are: one model line, the company itself, and no.
Investor note: the prevailing story assumes that frontier companies agreeing to slow down will depress future demand for training compute. This evidence weakens it: the only public allocation reading shows that compute reined in on one model line flows to the lines next to it, and the total drops only about 2%.
What would prove this wrong: any frontier lab publishes compute allocation shares with the whole company as the denominator, split into training, external inference and internal R&D, signed off by named outside evaluators; or a research group estimates a lab's training-versus-inference split from public contracts, power or chip data alone, and company figures later confirm it; or a pacing agreement writes down a named speed or total-volume cap and is carried out without any antitrust exemption. Verdict date: March 15, 2027, an observation window we set ourselves.
Why this matters to you: a cheap public rental price doesn't mean you can sign at it, and if you're sizing up a neocloud's debt, don't use its short-contract price to judge whether it can pay.
CoreWeave is the largest GPU rental cloud in the US, specializing in renting NVIDIA GPU compute to AI companies, and it's listed on Nasdaq. On September 17 it announced a $3B convertible bond due 2033. The investor presentation it filed with the SEC the same day says "~25% price increase across SKUs in July 2026," where SKU means GPU model. The same page says short contracts of 3 to 6 months signed in the third quarter recently priced at about $40M per MW per year, and that about 70% of contracts signed in the second quarter included customer prepayments (CoreWeave investor presentation, SEC 8-K exhibit, 2026-09-17). Annualized revenue per MW is this business's unit economics: how much a year you can collect from the GPUs that one megawatt of power can run. A footnote in the deck says the figure is those short contracts' annualized revenue divided by the power needed for the clusters that serve them.
The comparison is public pay-as-you-go GPU pricing, the list price for renting by the hour with no contract. getdeploying, a third-party site that tracks GPU cloud price lists, shows its index up 3.3% over the four weeks to the week of September 14 and up 5% over the past year (getdeploying, 2026-09). A second index, updated September 16 by the tech research and comparison site AIMultiple, puts the median pay-as-you-go price of the H100, NVIDIA's workhorse data-center GPU, at $3.25 per GPU-hour (AIMultiple, 2026-09-16).
Verification: we took CoreWeave's numbers straight from the SEC filing and checked the figures and labels on page 7 against the original chart. ⚠️ This is a marketing document the issuer handed over while raising money. It's "furnished," which carries lighter legal liability than a quarterly report, and the company has every reason to make its unit economics look good. The 25% is an adjustment to its own list prices, not an average market transaction price, and the deck doesn't say whether it applies to new contracts, renewals or everything. ⚠️ For the public index's rise we have only getdeploying so far; AIMultiple gives a level with no change over the same period, so it can't serve as a second source. ⚠️ The load-bearing seller-side numbers in this item come from CoreWeave alone.
Judgment update: using second-quarter figures from the same deck, here's the rough math: revenue of $2.575B times four, divided by roughly 1.5 GW of active power, gives a fleet-wide average of about $6.9M per MW per year, under a fifth of the short-contract price. The two numbers use different numerators and denominators, so read them only as orders of magnitude. On that basis we're taking on a working read we haven't settled yet: GPU compute really is tight this fall, but the premium sits in a thin slice of newly signed, short-term contracts. Neither the public list-price index nor the fleet average can see it, and by putting that slice's price in a bond deck, the seller is asking creditors to picture the whole fleet at its most expensive sliver. That runs opposite to a single-source reading we took in July, which read a falling spot hourly price for one GPU model as a sign that compute supply had overtaken demand. Both can be true at once. We're keeping the contradiction on the record and not rewriting the old reading.
Investor note: the prevailing story assumes public GPU rental prices work as a thermometer for the compute cycle. The CoreWeave deck cuts against that: if the premium lives only in new short contracts, using public spot prices to call a glut will systematically call it too early. For teams negotiating compute contracts now: if a seller anchors its quote on new short-contract prices, put long-contract prices and public pay-as-you-go prices side by side when you compare.
What would prove this wrong: CoreWeave's third-quarter report, due around November, shows revenue divided by active power rising sharply toward the short-contract price. The same would hold if the public index posts a quarterly rise close to double digits in the fourth quarter this year. GPU-only rental clouds such as Nebius and Lambda disclosing flat or falling new-contract prices over the same period would also count against us. Verdict date: December 15, 2026, set by us.
Why this matters to you: cloud providers' profits ride on how many years an AI chip stays useful. This is one of the few public readings pointing the other way, but it's just two examples the seller picked.
Read this alongside item 2 of today's main line. Page 10 of the same CoreWeave deck lists two cases from unnamed customers. One extended its renewal of NVIDIA A100s, first sold in 2020, through 2029, a "3-year contract signed in line with typical 1-year terms," meaning no discount for the longer commitment. The other renewed H200s, first sold in 2023, for three years at a "Premium to original contract" (CoreWeave investor presentation, 2026-09-17). Running A100s until 2029 means a card in service for about nine years.
Verification: we read the original filing directly. ⚠️ These are two cases the seller chose for a fundraising document, not fleet statistics; the deck gives no size for the H200 premium, no overall CoreWeave renewal rate and no resale value for older cards. ⚠️ On public indexes, the median pay-as-you-go price for an A100 is $1.76 an hour at AIMultiple (AIMultiple, 2026-09-16) and $2.00 on getdeploying's GPU cloud dashboard (getdeploying, 2026-09); both give only price levels, not renewal rates.
Judgment update: one short thesis in the market holds that the economic life of AI accelerators is far shorter than the server depreciation schedules on the big cloud providers' books, so reported depreciation is systematically too low. CoreWeave's two cases point the other way: a card nine years into service can still land a three-year contract. That weakens the idea that AI chips are generally useless within a few years; it doesn't rule out a gap between book life and economic life. What would actually settle economic life is renewal rates and resale values, and nobody has published those.
Investor note: the prevailing story assumes AI chips have short economic lives and cloud providers under-depreciate them on their books. This evidence weakens it, but only lightly: it's two examples the seller picked, with no renewal rates or resale values.
What would prove this wrong: CoreWeave or another GPU rental provider discloses a low overall renewal rate for older GPUs, or resale values for used A100s and H100s show up publicly and have fallen sharply. Verdict date: December 15, 2026, the same day we recheck item 2 of today's main line.
What to take away today: the only public allocation reading shows that compute reined in on one model line flows to its neighbors, and the total drops only about 2%. And if the premium lives only in new short contracts, reading a glut off public spot prices will call it too early.
1. [This month] (September 2026; we couldn't find the episode's exact release date) Researcher Noam Brown said on the show hosted by tech interviewer Dwarkesh Patel that once AI does its own research, progress could move 50% faster, three times faster, or ten times faster (Dwarkesh Patel, 2026-09). ⚠️ We haven't finished the full transcript. With a range that wide, if a capability curve became the acceptance test for slowing down, both sides could claim they were right.
No chips & semiconductors item this issue. This week's most important chip-side readings are CoreWeave's self-reported GPU prices and older-card renewals, covered in items 2 and 3 of today's main line.
*[This month] (posted September 9, look-back) Sebastian Raschka, a machine-learning researcher and author of Build a Large Language Model (From Scratch), warns that when you compare models on coding, the harness itself plays favorites. The harness is the layer of software wrapped around a model that decides, step by step, which files it reads, how much history it carries and when it calls tools; Anthropic's Claude Code and OpenAI's Codex are both harnesses. He writes: "models are typically developed with one primary harness in mind (and fine-tuned less on other harnesses). Plus, the primary harness is often developed to suit and amplify a model's strengths." (Sebastian Raschka, 2026-09-09). Which of our judgments it backs:* part of the cost gap between two labs' models on a third-party leaderboard comes from the harness, not only the model (the leaderboard numbers are in item 2 of Model watch below).
1. [This week] (the study priced at September 1 list prices; we read it September 18) [Measurement] UC Berkeley's HarnessTax study: the same model in a different harness can cost up to 5.1x as much per solved task. The study paired 7 models with different harnesses for 21 combinations in all, and ran each on two benchmarks, code repair and terminal operation, for 30 tasks each, 3 runs per task. Between the cheapest and most expensive harness for the same model, cost per task differed by 1.1x to 5.1x. The easiest example to remember: moving OpenAI's GPT-5.6 Sol from Claude Code to the Pi harness cut the cost per solved task from $1.540 to $0.441 (as relayed by Tomasz Tunguz, partner at venture firm Theory Ventures, 2026-09; study project page). Pi is another harness the study tested; the source doesn't say who makes it. ⚠️ Equal accuracy hasn't been shown. Across 42 comparisons under a strict statistical test, none had a gap big enough to rule out chance, but with 90 runs per cell the test can only detect gaps of about 15 percentage points or more. The researchers themselves write that this sample can't prove accuracy is the same, and can't say there's no gap either (in their words, "undemonstrated at this sample size, not proven absent"). ⚠️ We didn't read the paper itself; we checked the numbers two ways, through Tunguz's account and search-result summaries, and the two may share a source. Judgment update: the first reading in this direction was Raschka's personal test in June, where Claude Code burned 578,000 input tokens on one task and Codex used about half as many (Sebastian Raschka, 2026-06-27). Today brings a second, independent measurement, so we're moving "the unit cost of AI coding is set mainly by how the harness manages context, not by the model's price per token" up to a judgment still being verified, not yet settled: only two coding benchmarks were measured, and a quality gap hasn't been ruled out. For teams evaluating or buying AI coding tools: when you compare cost, treat the harness as a separate variable and hold it constant. Swap only the model without looking at the harness, and the gap you measure may be the harness's doing.
2. [This week] (third-party leaderboard read September 18) OpenAI says its new flagship GPT-6 Astra is about 63% cheaper per task than Anthropic's Claude Fable 5.1. On a third-party leaderboard the two tie on score and Astra's total spend is about 47% lower, but it's about 32% higher than OpenAI's own previous model, Sol. The receipts say: contested (the scores match; the cost gap is inflated). OpenAI's September 9 enterprise launch post puts Astra at 57.9% and Fable 5.1 at 55.8% on Terminal-Bench 4.0, a benchmark of multi-step terminal tasks, and claims "approximately 9% and 63% lower estimated API cost per task" versus Sol and Fable 5.1 (OpenAI, 2026-09-09). On the third-party leaderboard maintained by Snorkel AI, a data-labeling and AI evaluation company, at maximum effort Astra scores 58.2% and Fable 5.1 57.9%, a tie within the margin of error. Total spend to run all 66 tasks: Astra $3,300, Fable 5.1 $6,200, Sol $2,500 (Snorkel AI, read 2026-09-18). ⚠️ Different measures: the vendor post gives an estimated per-task cost with an undisclosed method; the leaderboard gives totals, not normalized to a common basis by effort level. ⚠️ Each model runs in its own lab's harness, so per item 1 there's no telling yet how much of the cost gap is the model and how much the harness. ⚠️ Disclosure: our research system runs on Anthropic's models; this item only relays the leaderboard and the launch post.
[This week] (launched September 17) OpenAI packaged its flagship model as "Astra for Law," a legal-industry product, walking straight into the application layer held by legal AI startups such as Harvey and Legora. It connects GPT-6 Astra to a US legal research index covering more than 230M URLs, plus legal-writing instructions and privacy controls. Selected law firms get it first, with an API coming soon; Harvey and Legora can build products on it while also competing with it. On 200 private legal-research questions from Vals AI, an AI evaluation platform, OpenAI's own run scored the legal version at 54.0% and the same model with only web search at 38.7% (OpenAI, 2026-09-17). ⚠️ This is OpenAI's self-report, comparing its own product with itself, and we have no second source. Our observation: it's the same kind of effect as item 1 of Model watch; give one model a different retrieval and instruction setup and its performance shifts a notch. Our inference: for any team whose product is "model plus domain retrieval plus instructions," this is the same risk. The model maker can build that layer too, and holding your ground depends on what it doesn't have, such as workflow integration and existing customer relationships.
No archive pick this issue. The older material we could use has run out.
The past 24 hours. September 17 to 18 added 236 pieces: 121 papers, 50 paper summaries, 43 blog posts, 13 podcast transcripts, 7 subscription newsletter issues, 1 industry analysis and 1 company filing. Of those 236, last night we finished reading 2 and set aside 1, leaving 233 unread; papers and paper summaries together make up 72% of the new material. What we read was CoreWeave's filing; the one set aside was a blog post relaying something we already wrote about on September 17. All 13 podcast transcripts are older episodes of a single show, backfilled at once. Nothing new came in from social platforms last night. Of the 17 outside receipts used in the body, 3 come from last night's 236 pieces: Zvi, Dwarkesh Patel and the CoreWeave filing (for the Zvi and Dwarkesh pieces we read only the passages the deep-dive column cites). The other 15 we fetched from their original addresses or pulled from what we already hold.
Older material in this issue. No sources were added in a one-time backfill today. The September 9 post in Named commentary and the June 27 test cited in Model watch are older material we already held, not part of the past 24 hours.
Source concentration. No single source accounts for more than a third of the load-bearing sources. But the seller-side numbers in items 2 and 3 of today's main line all come from CoreWeave alone, and all from a fundraising document; the study numbers in item 1 of Model watch all come through a single relayed account.
What you are not getting today. One: we didn't read the HarnessTax paper itself; its project page loads dynamically. Two: as of this issue we hadn't found the pricing outcome of CoreWeave's convertible bond. Three: we didn't read the original text of Woodward's remarks at the Justice Department or of Whitesides's op-ed. Four: one GPU rental-price report blocked our fetch, so the public prices in item 2 rest on two indexes only.
The sources we track. 529 named speakers in total. The spread: social platforms 302, podcasts 90, outlets 51, blogs 48, paper authors 48, newsletters 46, earnings calls 26, keynotes 23 and a scattering of others; another 77 company and institutional blogs aren't counted. ⚠️ Those count venues, and one person can appear in several, so the parts add up to more than 529. Representative names: in newsletters, Zvi Mowshowitz, Gary Marcus, Dwarkesh Patel and Sebastian Raschka; in blogs, Tomasz Tunguz; on the institutional side, SemiAnalysis and the NVIDIA Technical Blog. Several identically named numbers count different things. The roster's 46 newsletters, 48 blogs, 90 podcasts and 48 paper authors count speakers we track over the long term; last night's 7 newsletter issues, 43 blog posts, 13 podcast transcripts and 121 papers count articles that came in over one night. Different populations. This issue uses 17 outside sources in the body, the same figure printed in the footer, counting only links the body actually cites that are not on our own domain; that is also a different scope from last night's 236 pieces.
This is not a news digest: we hunt each day's AI firehose for the insights that actually matter and the practitioner judgments worth tracking over time, and we show how every item was verified — the point is always "which judgment got harder, and who's been right," never "what happened today."
— SecondSource · generated by our research system · 17 sources · Got a view? Reply and tell us
Written from the same research and judgments as the Traditional Chinese edition; every claim links to a primary document.