From the Aug 2, 2026 daily brief
First, the background. On February 3 this year, US software stocks suffered a historic one-day crash, triggered by several AI labs shipping agents that can complete long tasks autonomously, all in the same week; brokerage traders dubbed it the "SaaSpocalypse." The next day, the Taiwanese investing podcast Gooaye proposed a reframe in episode 633: stop asking "will AI kill software" and start asking "can this company pass its token costs through to customers." AI gave software a real variable cost for the first time: every model call costs money, metered in tokens. Tokens are the units a model slices text into for billing; they are not words, and the same text can produce more or fewer of them depending on how it is sliced. Companies that can pass the cost through see their gross margins expand, the argument went, because their products are strong enough to charge usage fees on top of seat fees; companies that cannot must swallow it, because their products are weak enough that a competitor can knock off a substitute with AI. Those are the ones that get killed. This newsletter first logged that framing as a hypothesis to verify while working through a backlog of Gooaye episodes. Six months after episode 633, this essay checks the intervening public filings, official price sheets, and executive testimony against it line by line. Three conclusions came out of the audit:
>
- The direction (software splits into two camps) has been confirmed by the market more forcefully than anyone expected. Inside the same software index, the strongest and weakest constituents of 2026 sit roughly 140 percentage points apart year to date (Datadog up more than 80%, Atlassian down more than 60%); Goldman Sachs went ahead and built a buy-this-side, short-that-side pair-trade basket. The index itself is nearly flat on the half year while individual names bleed out. The split itself is now a fact.
- But "can you pass it through" is not the dividing line, because in eighteen months the pass-through plumbing became a commodity anyone can buy. The five bellwethers this essay tracks — Notion (notes and collaboration), Canva (design tools), Zoom (video conferencing), GitHub (code hosting), Figma (interface design) — started from opposite ends and converged on one billing architecture: bundle the AI into the subscription, set a usage floor, and meter whatever overflows. Industry surveys point the same way: within a year, companies giving AI away free fell from 34% to 15%, while usage-based pricing rose from 19% to 42%. When everyone can pass costs through, pass-through no longer separates the living from the dead. What actually separates companies is two deeper variables: whether a company's AI consumption is bounded (one click, one generation — or an agent left running all night), and whether its workload can downshift (move from the priciest flagship model to a cheaper tier that is good enough).
- The most original half of the hypothesis — that successful pass-through companies see margins expand — came out half right, half wrong. Margin expanders genuinely exist (Palantir went from 80% to 87%, Salesforce from 75.5% to 77.7%, Duolingo recovered two points in a single quarter), but nowhere in the filings does any company attribute the expansion to passing costs on to customers; the only company that states a cause, Duolingo, credits "continued reductions in per-unit AI costs and improved efficiency." The real margin path is swallow first (Notion's CEO, asked whether he would let gross margins get worse, answered "Oh, you have to"), optimize second (Zoom gave its AI away free, watched margins dip, and two years later recovered to above where it started), and only then recover. The life-or-death line is not whether you can hand the cost to your customers: it is whether, after swallowing it, you can optimize it back out.
Start with the hypothesis in its own words. Gooaye's two scripts run like this. Scenario A: you used to sell seats (the traditional enterprise-software model of a monthly fee per user), where one finished product sells to a hundred thousand customers at near-zero marginal cost; now there is AI in the product, every customer use of it costs you a model fee, and your gross margin gets squeezed. Scenario B is the same event read in reverse: AI supercharges your product, so you can charge a metered token fee on top of the old seat fee — while the model costs you pay to cloud providers behave, in the host's words, "just like cloud costs … it becomes a kind of infrastructure, like a utility bill," and fall over time (our translation from the Mandarin). Revenue rises with usage, costs fall with time, and margins expand instead. His judgment was that both scripts play out simultaneously at different companies, and the sorting variable is product strength: "some companies will only ever get Scenario A, because their product just isn't competitive enough … someone can vibe-code something that competes reasonably well" (vibe coding: building software conversationally with an AI, no traditional engineering team required).
The hypothesis deserves serious treatment because its first half stands on a line of evidence this newsletter has checked many times. Microsoft CEO Satya Nadella and his CFO formally recast Microsoft 365's business model in their earnings language from "seats" to "seats plus consumption," and Ben Thompson's read is that this was not voluntary innovation: agents forced it. An agent does more work than any single person, consumes far more compute than a person, and shrinks the number of seats a company needs; the per-seat foundation is knocked out, and software, carrying a real marginal cost for the first time, will eventually pass it through. We hold that judgment only as a relay, with no linkable primary source, so it is treated as single-source. But the same line found independent, cross-industry confirmation in the first half of 2026. The CEO of chip-design-software giant Cadence described his own agent pricing more bluntly: traditional licenses keep selling, plus a new per-agent layer priced "based roughly on the amount of work one human might do," plus a pure consumption layer like a model vendor's. He reported his customers' attitude in so many words: "They would rather pay for that than hire more people" (More Than Moore). Office software and chip-design software, two industries that do not cite each other, grew the same three-layer structure.
One timing fact needs stating up front. It does not diminish the hypothesis, but it changes how you grade it: episode 633 was recorded the day after the software crash, so it is an attribution framework for something that had already happened, not an advance prediction. What this essay audits is not "did he call the crash"; it is whether the sorting mechanism he proposed and the sorting the following six months actually produced are the same thing.
The hypothesis quietly assumes that pass-through is a scarce capability — some companies can, some cannot. Six months on, that assumption died first.
Start with the industry-wide picture. A tracking survey of roughly three hundred software companies shows that within one year, the share giving AI away free fell from 34% to 15%, consumption pricing rose from 19% to 42%, and outcome-based pricing rose from 2% to 23%; asked why they repriced, the second-most-common reason (34%) was margin erosion (ICONIQ, State of AI 2026 — we checked the underlying PDFs page by page). An annual B2B pricing survey points the same direction: pure seat pricing fell from 21% to 15% in a year, and hybrid pricing (seats plus usage) rose from 27% to 41% (Growth Unhinged). Even the billing plumbing is pre-laid: Stripe's Token Billing lets any AI application peg its retail price to upstream model costs in real time, formula printed right in the product docs (Latent Space). Everything pass-through requires — pricing design, metering, billing infrastructure — is now available off the shelf.
More persuasive still is the convergence of the five bellwethers. Notion's AI went from a $10-per-user-per-month add-on in 2023, to "unlimited" inclusion in the Business plan in May 2025, to charging credits for custom agents in May 2026 ($10 per 1,000 credits, verifiable on the official pricing page). Canva walked the same road backwards: it metered credits from the start, then in April 2026 stacked an almost-all-you-can-eat tier on top at $100 a month (Fortune). Zoom gives AI Companion away free to every paid seat and sells a customized tier at $12 per user per month. GitHub Copilot went furthest of all. From June 1, 2026, the entire product line moved to usage-based billing: plan prices unchanged, an equivalent allotment of credits included, overages deducted at model list prices. The GitHub product executive who announced it wrote the reason on the official blog himself: "Today, a quick chat question and a multi-hour autonomous coding session can cost the user the same amount" (GitHub Blog). That sentence is an official admission: in the agent era, flat-rate pricing cannot cover the cost. The fifth company, Figma, has the most complete story of the five; it waits for audit two.
Why did everyone converge on the same architecture from opposite starting points? Because agents stretched the usage distribution into an extreme skew. OpenAI's CFO has disclosed that paying heavy users run eleven times the usage frequency of free active users; a human typing by hand burns under 100,000 tokens a day, an agent starts at 100 million a day, and one extreme user burned 130 billion tokens in a single month (Exponential View). Subscriptions presuppose a natural ceiling on usage; agents have no physiology. So every company eventually grew the same shape: bundle (keep the subscription budgetable) + a usage floor (block the tail) + metered overflow (make the outliers pay for themselves). Notion and Canva both decline to publish where the floor sits. That is no coincidence: they are two solutions to the same arithmetic.
By August 2026, "can you pass it through" is the wrong question. Everyone is passing through. The only suspense left is whether customers walk after you do, and that is the old question of product strength, a full circle back to differentiation, not a new dividing line.
Scenario B's cost-side premise is that token prices fall over time like a utility bill. This newsletter has verified that claim repeatedly on the seller side (the model vendors), and for this audit we re-checked the official price sheets from late 2025 to now, sheet by sheet. Conclusion: "falling" and "rising" are both true, depending on which tier of the sheet you stand on.
Look down, and the low tiers are collapsing. On July 30, OpenAI cut the low-end model of its new flagship family by 80% (input price slashed from $1 to $0.20 per million tokens), confirming what The Wall Street Journal reported in June, that OpenAI was weighing drastic cuts (Marcus on AI), with the blade falling entirely on the cheap tiers. Google's lightweight model dropped 80% in a single generation. The academic measurement is more systematic still: fix the capability level (say, "matches GPT-4 on PhD-level science questions"), and the price of acquiring that fixed capability falls between 9x and 900x per year, median 50x (Epoch AI). On the "same task" yardstick, the utility-bill thesis holds completely; the deflation is startling.
Look up, and the price of admission to peak capability is rising, in four distinct forms. First, outright flagship repricing: OpenAI's flagship line raised input prices within a year to four times their starting level ($1.25 → $5). Second, Anthropic opened a new tier above the flagship: Fable 5 sits at twice Opus, with fast mode charging a further 2x premium on top. Third, discounts are expiring on a published clock: the docs state Sonnet 5's current discount ends August 31, with standard pricing back in force September 1 — a 50% increase. Fourth — and the least noticed item in this audit: Anthropic's official documentation states that the new generation of tokenizer "produces approximately 30% more tokens for the same text" (Anthropic pricing docs). List prices unmoved; the bill for the same document up 30%; invisible in any price-sheet comparison. While we are at it, one claim this newsletter previously logged gets corrected on the spot. The Glean CEO's July line that every model vendor had raised per-token prices over the past six to nine months (20VC) fails the direct check: in that same window Anthropic cut its mainline Opus tier by 67%. But the thing he was pointing at is real; it just had the wrong name. What rose is not "everyone's prices." It is the price of admission to the current frontier.
This tiered price sheet matters more to software companies than to model vendors. It rewrites Scenario B's cost premise into a conditional: whether your costs fall like a utility bill depends on whether your workload can downshift, moving off the flagship to a cheaper tier that is good enough. Those who can move ride the 50x-a-year deflation; those who cannot face repricing, new premium tiers, and invisible inflation all at once. And whether you can move is a matter of first-hand testimony. Notion CEO Ivan Zhao drew the line precisely in a May interview. Coding products have to run frontier models, usually the smarter the better, but knowledge work can already settle for second-tier open-weight models: "a lot of this kind of paper-pushing information in your company doesn't require Opus to file tickets" (Sequoia, Long Strange Trip). The enterprise version of the same judgment comes from Glean's CEO: ninety percent or more of enterprise use cases today no longer require the most powerful models (as relayed; single source).
Open ?
Here is the one sentence that frames how to read this tree: it is this newsletter's main battleground for tracking whether model vendors get to keep the margin on the tokens they sell, and until now we have always read it from the seller's side. This essay adds the buyer's mirror image: the sellers' tiered pricing, landing on buyers, becomes a wealth gap in downshift freedom, and that gap is replacing "can you pass it through" as the real dividing line on software's cost side.
Now for the most consequential half: Scenario B says the strong companies' margins are "not squeezed — they expand." This essay drew on three lines of evidence: raw gross-margin figures from SEC filings (computed ourselves, not relayed), full earnings reports and call transcripts, and executive interviews at the private companies. The verdict comes in four layers.
Layer one: swallowing the cost is the universal transition state, unrelated to product strength. Two founders whose product strength nobody disputes speak first. Notion's Zhao was asked point-blank by a HubSpot co-founder what is happening to his gross margins and where they settle out; his answer: "I don't have an answer for you." Pressed on how far he was willing to let margins degrade: "Oh, you have to." Canva co-founder Cliff Obrecht had already put a number on it. Asked whether AI costs had reached a tenth of revenue, he answered: "100% yes, it already is. If you look at Lovable, their pass-through to Anthropic or model providers will be way more than 10%" (SaaStr, relaying a 20VC conversation; a co-host's firsthand write-up, not checked against the original audio). The largest company of all reads the same: Microsoft's latest annual report shows company-wide gross margin falling two years running (69.8% → 68.8% → 67.9%), attributed in the filing to "continued investments in AI infrastructure and growing AI product usage." For the first time, inference costs failing to decline appears among the named risk factors (Microsoft FY2026 10-K).
Layer two: Figma acted out the hypothesis's entire mechanism, filings included. Figma began by giving AI even to free users and swallowing all of it: gross margin collapsed from 88.3% in 2024 to 82.4% in 2025, with filings attributing the surge in cost of revenue chiefly to AI inference and hosting (up another 253% in Q1 2026, AI the largest component). Then it turned: from March 2026 it sold AI credits ($0.03 apiece) and made per-seat credit limits mandatory across the board. Its CFO, in prepared remarks: "With full seat AI credit limits now live, growing AI usage and adoption now translates into revenue — a key monetization milestone," listing the cost-side levers as routing queries across models by task complexity, a model-agnostic architecture, and investment in first-party models (Figma Q1 2026). Worth flagging above all is the risk factor in its 10-Q, which in effect writes Gooaye's Scenario A into a legal document: "Pricing pressure may require us to … absorb additional costs, or operate certain AI features at lower or negative gross margins for extended periods … competitive dynamics may result in customers expecting full AI functionality to be included in baseline subscription pricing." One company, twelve months, from Scenario A to Scenario B. The turning point was installing the meter.
Layer three: margin expanders exist, but not one attributes the expansion to pass-through. This is the most honest ruling on Scenario B the evidence allows. The expansion cases are real: Palantir's gross margin walked from 80.2% in 2024 to 82.4% in 2025 to 86.8% in Q1 2026, expanding and accelerating through the AI era; Salesforce went from 75.5% to 77.7% across the entire Agentforce rollout; Atlassian expanded from 82.8% to 84.2% while giving AI credits away free. Zoom gave AI Companion away, took the dip (its annual report states the decrease was "mainly due to increased costs associated with the use of AI functionality"), then recovered; its latest quarter's 77.9% is its highest since 2022, with the CFO stating the mechanism plainly: "as the AI costs spike in a good way with usage, we have got offsetting measures." And the only company in the entire field to put a causal sentence about rising margins into a filing is Duolingo: Q1 2026 gross margin recovered from 71.1% to 73.0%, and the company attributes it to "continued reductions in per-unit AI costs and improved efficiency." This is the same company that a year earlier wrote that margins fell on "increased AI costs used in features like Video Call." Its CEO's description of the cost curve: "Costs are coming down. They've come down just without us doing anything." Put these together and the engine of expansion is unambiguous: not handing costs to customers, but audit one's same-capability deflation curve (50x a year) flowing through to the buyer's side, plus model routing, plus the discipline of locking AI inside the paid tier (Duolingo's principle: AI features live in the Max subscription, where "the usage of AI is anyways profitable"). Industry-level data agrees: average AI-product gross margins went from 41% in 2024 to 45% in 2025, the latest report says they are "on track for ~59% by 2027," and two-thirds of companies report per-query costs improving (ICONIQ, cited above; the forward number is a projection, not a measurement). The frontline analyst's winning condition is accordingly the opposite of Gooaye's. Altimeter's Jamin Ball, writing in March: "The ones who figure out that abstraction, pricing on value and managing token costs as an internal optimization problem rather than a customer-facing one, will build the most durable businesses" (Clouded Judgement).
Layer four: the cost variable that actually cleaves companies in two is whether usage has a bound. Same period, same model prices; two groups of companies sit 100 gross-margin points apart. The bounded group — one click, one generation, budgetable per call (Zoom's meeting summaries, Duolingo's lessons, Adobe's image credits) — runs margins between 70% and 89%; Adobe's overall margin has barely moved (around 89%; look closely inside the subscription line and there is very slight compression: subscription costs growing a bit over two points faster than subscription revenue, the highest-margin Digital Media segment slipping from 96% to 95%, magnitudes of a few tenths of a point, with "AI inferencing costs" listed explicitly in the cost definition). The unbounded group — coding tools running agents on long tasks — floats, according to relayed reports, between negative and forty percent (Cursor around negative twenty to thirty, Replit swinging between 36% and negative 14%, Windsurf "very negative"; all media relays, unaudited, treated as single-source). Cursor's own apology post: "the hardest requests cost an order of magnitude more than simple ones." One caveat: this sample concentrates in agent-running coding tools. But Figma, which is not a coding tool, crashed its margins during exactly the stretch its free tier ran unbounded, so the operative variable looks like boundedness, not industry. And note this group's products are, by common consent, the strongest in the field. Predicting margins from product strength fails outright on this data; predicting from boundedness of consumption scores a clean sweep. It also explains pass-through behavior itself: the unbounded companies were forced to pass through earliest (GitHub, Cursor, Replit all moved to usage billing in 2025–2026), while the bounded ones can afford to give AI away. Pass-through is the effect, not the cause.
Finally, the "margins haven't moved at all" evidence belongs in the picture too, because it is simultaneously true: a benchmark study of 342 companies finds the software industry's median gross margin steady at 79–81% across four years, concluding in so many words that "AI infrastructure costs have not yet compressed software margin at the median" (Benchmarkit × Aleph). This contradicts nothing above: today's compression and divergence live at the AI-product-line layer (margins around 45%) and have not yet propagated to the company layer (around 80%). But AI products' share of software-company revenue, per the same tracking, is projected to pass half in 2027. The two lines are on a collision course, and the impact point is 2026–2027. The same study carries a cross-sectional warning: companies on usage-only pricing show a median margin of just 62%, far below the overall 80%, meaning "moved to metered billing" and "high margin" are negatively correlated in cross-section (because the companies that have moved are mostly the cost-heaviest AI natives). "Margins improved after metering" holds only in longitudinal, same-company narratives (Figma, Duolingo). Reading the cross-section as if it were longitudinal is the easiest statistical mistake in this whole topic.
Last, the market half: the hypothesis predicts software stocks split into two camps by pass-through ability. The split arrived, with startling force. The software index traced a V and gave it back in 2026 (January was its worst single month since 2008; the April trough hit minus 30% from the start of the year; June briefly went positive; late July fell back to about minus 12%). The index went nearly nowhere in half a year while roughly 140 percentage points opened up between its constituents: Datadog up more than 80% and Snowflake up 35% in a single earnings day on one side; Atlassian down more than 60%, and Salesforce and Workday each down around 30%, on the other. This matches the sector dispersion this newsletter logged in July: five software categories with similar growth rates and wildly different returns, the market rewarding the categories AI-era CIOs deem essential and punishing per-seat horizontal apps (Tomasz Tunguz).
But look at the criterion, and the market is not sorting on pass-through. Goldman's software long/short pairs trade goes long on companies that need physical infrastructure, carry deep regulatory entrenchment, or demand human accountability, and shorts names built on seat-based pricing or workflow automation (as relayed; single source). Cost pass-through appears nowhere in it. Set the two knives side by side: Gooaye's knife asks "can your cost side take the hit"; the market's knife asks "will your demand side be replaced by AI outright." The two correlate but do not coincide, and in the two named cases where their predictions fight, the market's knife won:
Salesforce, which by the pass-through thesis should be thriving, got cut anyway. It is pass-through's honor student: Agentforce bills by usage, its annualized revenue blew past $1.2 billion from $800 million within a year, gross margin kept expanding across the rollout, and core revenue growth accelerated from 8–9% to 13% (official filings). Pass-through succeeded, fundamentals accelerated; the stock still fell thirty to forty percent on the year. What the market is cutting is the endgame narrative: whether agents pull the floor out from under the per-seat core business. Palantir, which by the pass-through thesis should be camp B's flag-bearer, is falling too. Gooaye named it in the original episode as the example of camp B "already priced in early," and its fundamentals have played the Scenario B script to the hilt (margins out to 86.8%, revenue guidance up about 70%), yet the stock trails the market, down about 20% on the year. Belonging to the winners' camp guarantees no return in this window; valuation digestion dominates everything. Together the two cases say: the sorting framework can explain income statements; it cannot pick stocks. And a tradeable sort is exactly what the original thesis was for.
Palantir deserves two more sentences, because it is habitually cited as the pass-through exemplar and the evidence is more complicated than the meme. Its pricing is actually discoverable, and it is not seat-based: UK government procurement pricing documents show an annual organization fee with an equal amount of usage included and "No additional user licences required," all computation metered in compute-seconds, and model "tokens are converted into compute-seconds to match the price of the underlying model provider," drawn down against the prepaid pool (Palantir docs). In plain language: the enterprise pays one annual fee that already includes an allowance of computation; every kind of computation converts into a common unit, compute-seconds, for metering; model calls are charged at the model vendors' own list prices out of the prepayment. Formally, that is indeed usage pricing. But reading its margin expansion as a pass-through victory has a checkable hole: the top line item in its cost of revenue is people, with third-party cloud hosting fifth; software and services margins are never disclosed separately; and Q1 revenue surged 85% against costs up just 25%. The bulk of the expansion looks like operating leverage in the services organization plus customer-mix shift. The honest reading is an asymmetry: the companies that write AI costs into their margin attributions (Microsoft, Adobe's subscription line) are all compressing, while the company whose margin expanded most happens not to break out where the margin comes from. (Palantir reports Q2 after the close on August 3, the nearest verdict window after this essay goes to press.)
Six months of evidence narrows "camp A gets killed by vibe-coded copies" to a stratum effect. In the small-and-mid-size tier it is real: named cases of 20-to-70-person companies dropping Salesforce and HubSpot contracts, building replacements with AI tools, and cutting software costs 40–80% now arrive in batches (PYMNTS, relaying The Information); Retool's survey of 817 teams says 35% have already replaced at least one SaaS tool with something self-built (Retool). Go upmarket and the evidence thins fast. At the large-enterprise tier there is exactly one case so far, Sanofi, which cut ServiceNow usage by 80% rather than cancel (Fortune); and the most-cited flagship case, Klarna "replacing Salesforce and Workday with AI self-builds," has been debunked: it swapped in a different batch of SaaS and remains a Salesforce partner (CX Today). Complex business logic is still the ceiling: even Wix, which sells vibe coding, demonstrated that two teams of professional engineers could not build a barbershop-management system with its own tools (20VC). Gartner's official yardstick: by 2030, about $234 billion — roughly a fifth of enterprise application software spend — is "at risk" (Gartner). That is the magnitude of an erosion, and it comes with four years attached.
The software split is real, but the line is not "can you pass token costs through." The pass-through plumbing commoditized within eighteen months: five bellwethers converged from opposite starting points on the same bundle + usage floor + metered overflow billing architecture, and companies giving AI away free shrank from 34% to 15% in a year. Once everyone can pass through, pass-through stops sorting the living from the dead. Six months of evidence points to three variables with real predictive power. On the cost side: boundedness of consumption (one click, one generation — or an agent running all night; the only variable that explains the 100-point margin gap between contemporaneous groups, on data where product strength fails outright) and downshift freedom (whether workloads can move from the flagship model to a good-enough cheaper tier; movers ride the 50x-a-year same-capability deflation, non-movers face admission prices rising on every side). On the demand side: replicability (whether your business logic survives customers building their own with AI — the market's knife, the criterion in Goldman's pairs trade). The real margin path is swallow, optimize, then recover: swallowing is the universal transition state (Notion CEO Ivan Zhao's "Oh, you have to"; Microsoft down two years straight), the recovery engine is unit-cost deflation plus model routing (Duolingo is the only causal sentence in any filing), and "expansion because of pass-through" remains without a single documented case. The original hypothesis's directional call (software splits in two) stands, with the evidence still strengthening; its mechanism call (the line runs through pass-through, and margins expand because of it) is hereby revised.
>
What would prove this wrong (observation windows are this newsletter's own choosing except where noted; verdict dates fall when they close): One — any listed software company, in a filing or on an earnings call, explicitly attributes rising gross margins to passing AI usage costs through to customers (rather than to cost optimization): "pass-through → expansion" revives and this essay's "no documented case" ruling is void. Two — Zoom's or Duolingo's gross margin falls back below its trough and stays there (Zoom below 75.8%, Duolingo below 71%, on an annual basis): the signature swallow-it-and-optimize-out arc has flipped over. Three — a genuine price war opens at the flagship tier (OpenAI or Anthropic cuts a current flagship by 50% or more, with no new higher tier opened and no tokenizer-style hidden clawback): the "rising price of admission" leg breaks, downshift freedom loses value, and the original utility-bill thesis revives. Four — any Fortune 500 company publicly cancels (not merely cuts usage of) a core SaaS product in favor of an AI self-build: large-enterprise replicability upgrades from one case to a trend, and the kill radius of the weak-product camp — the companies whose product a competitor can knock off with AI — must be revised upward. Five — the company-layer margin median (Benchmarkit's yardstick, currently 79–81%) breaks below 78% in the next annual read: the "compression is still confined to the AI product line" layering fails and the across-the-board-compression thesis gains weight; conversely, if AI revenue share passes half as projected while the median holds, "swallowing is a transition state" upgrades from reading to conclusion.
Software CEOs and CFOs. This audit changes the to-do list from "design your pass-through" to three items. First, copy the industry's convergent billing architecture outright: bundle plus usage floor plus metered overflow. The five bellwethers have already made the mistakes for you. But all five sell to enterprises or professional users; consumer subscriptions carry stickier inertia, metered overflow triggers churn directly, and the one consumer case in this essay, Duolingo, locks AI into the Max paid tier rather than billing overages. The two templates are not interchangeable. Figma's lesson: the later the gate goes in, the more you swallow first. Second, design consumption to be bounded: whatever can be one-click-one-generation should not become an always-running agent, and any agent you do ship gets a credit cap. Duolingo's lock-AI-in-the-paid-tier and Figma's per-seat credit limits are two directly copyable templates; your margin is set by your consumption shape, not your product strength. Third, build downshift capability: split your workloads into what genuinely needs the frontier and what a second-tier model can carry (Notion's razor travels well: paper-pushing doesn't need Opus). A company without model routing is forfeiting a 50x-a-year cost deflation. Write promotion expiry dates like September 1 and tokenizer generation changes into your cost forecasts too, because your model bill can rise even while list prices hold still.
Commercial leads at model vendors and cloud providers. Your customers are learning to downshift and route. That is why the low tiers must be fought to the bone while the flagships can raise prices: the cheap tier's competitor is "the customer routes to an open-weight model themselves," while the flagship's only competitor is another flagship. The real battleground of tiered pricing is the middle, where the customers who want to downshift but haven't yet built routing live. The move is to make helping customers downshift a product in itself — routing middleware, downshift advisory, mid-tier plans priced as a share of savings — rather than waiting for that cohort to learn routing on their own and hoping to catch them on the way down.
Investors. Three named lessons: Salesforce proves that successful pass-through plus accelerating fundamentals cannot stop an endgame narrative from crushing the multiple; Palantir proves that margins expanding to 87% guarantee no return; the Cursor group proves that the strongest product in the field can coexist with negative margins. Use the sorting framework as an income-statement tool — judging whose margins repair first, by boundedness of consumption and downshift capability, not by the pricing page — and not as a stock picker. The most honest signal on the watchlist is the margin-attribution language on earnings calls: the quarter it flips from "investment in AI" to "efficiency gains" or falling "per-unit AI costs" is the quarter the arc bottoms (Zoom's flip came in mid-2025, Duolingo's in Q1 2026).
Enterprise software buyers. Your leverage is far larger than it was eighteen months ago: the SMB self-build cases hand you negotiating ammunition, and Sanofi hands you the cut-usage-without-canceling template. But keep Klarna's lesson beside them: claims of "replacing SaaS with AI" routinely land, in practice, as "switching to a cheaper SaaS." Before deciding, price the self-build's maintenance, compliance, and security costs in full; Wix's barbershop experiment is the honest gauge of where the complexity ceiling sits.
Written from the same research and judgments as the Traditional Chinese edition; every claim links to a primary document.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The deep dive "The Guarantor Falls, the Guaranteed Rise: Circular Financing Is Now Priced …
Start by laying out what happened on July 27. Before Monday's open, The Wall Street Journa…
The foundry/IDM framework in this essay borrows heavily from Ben Thompson's long-running s…
Start with the exact words. On the BG2 podcast in December 2025 — with Arvind Jain, CEO of…