SecondSourceAI Industry Insight · Full Archive

Deep Dive No One Gets to Be TSMC: Why Everyone in AI Is Becoming an Integrator · Jul 27, 2026

No One Gets to Be TSMC: Why Everyone in AI Is Becoming an Integrator

Key takeaways

The semiconductor industry runs on a famous act of self-denial: TSMC does contract manufacturing only, never designs its own chips, and that promise ("we will not compete with our customers") created the entire fabless design industry. Two days ago, when this newsletter audited the model-commoditization debate, it landed on a judgment: frontier AI labs do not occupy TSMC's position. They occupy the position of a company that sells you the process and sells its own finished products: what the chip industry calls an integrated device manufacturer, or IDM, competing directly with their own API customers. This essay treats that judgment as a dashboard awaiting verification and does three things. It starts by reading each of the four monitorable gauges the judgment shipped with: all four have moved, and the gauge for "labs internalize the next vertical application", whose twelve-month observation window opened only two days ago, pegs at full reading on day one, because the supporting evidence had already happened between January and May. Then the harder question of why the labs have no choice but to enter their customers' markets: the answer is arithmetic. Staying in the pure token-selling position is a textbook perfect-competition trap, and TSMC could afford to stay a pure foundry only because it holds a monopoly on advanced process technology, which the labs have no equivalent of. And it reports a revision to the original judgment: integration is not a one-way move by the labs; the application companies whose business is being raided are training their own models in return, and the whole industry is converging on an equilibrium where everyone is an integrator. The genuinely valuable neutral position has not disappeared; it has moved up to the distribution and governance layer, and the market has already started pricing it.

The self-denying pledge that built an industry

The foundry/IDM framework in this essay borrows heavily from Ben Thompson's long-running semiconductor analysis at Stratechery; what this newsletter adds is six months of AI-industry evidence layered on top. Start by being precise about what "pure-play foundry" means. When Morris Chang founded TSMC in 1987, the model he committed to was to manufacture other people's chips and never make products of his own. That rule, which looks like leaving money on the table, is actually the core of the whole model's value: because the foundry makes no chips of its own, customers dare to hand over their designs, the crown jewels. Thompson has drawn the causal line bluntly: TSMC's distinctiveness lies precisely in not designing its own chips, which freed designers, previously bound by the rule that every company must build its own fab, to start companies of their own. NVIDIA got going in 1993 on just $20 million and never built a fab; Qualcomm and Apple later walked the same road (Stratechery). TSMC still reaffirms the rule every year: its 2025 annual report again commits the company to never deviating from the pure-play foundry business model, describing it as proven, time and again, to be a win-win for TSMC and its customers (TSMC annual report).

The pledge won on the financials too. In 2022 came a historic crossover: the specialization camp (TSMC as pure foundry plus AMD as fabless designer) saw both companies' gross margins overtake Intel, the holdout still designing and manufacturing under one roof (Stratechery). Specialization beating vertical integration has been the semiconductor industry's main theme for two decades.

One common extension of this story needs correcting first, though, including in this newsletter's own earlier writing (the July 25 piece auditing LLM commoditization against foundry economics made the same move): holding up Samsung as the cautionary tale of "lose neutrality, lose the business." Samsung's foundry share is indeed small: TrendForce has TSMC taking 72.3% of global foundry revenue in the first quarter of 2026 against Samsung's 6.5% (TrendForce). But when we went back to check, the mainstream analyst explanation for Samsung's lag is technology, not trust: Samsung's 2-nanometer yields have reportedly been stuck around 55%, below even the roughly 60% threshold needed for stable mass production (TrendForce). The mechanism, whereby competing with your customers costs you customers, is real; TSMC built a market on the inverse promise. But Samsung's plight is mostly a yield story, and pinning all of it on lost neutrality overstates the evidence. This correction does not overturn the essay's main judgment; it swaps the argument's foundation for harder ground: the value of neutrality shows in what it created, not in how badly those who lack it are suffering.

Four gauges, all moving

Back to AI. The labs-are-integrated-manufacturers judgment shipped with four monitorable gauges. Here is how each one reads.

Gauge one: first-party applications as a share of lab revenue. The benchmark number here got an upgrade. Claude Code's $2.5 billion annualized revenue was originally just a market estimate that surfaced after the supply-chain leak episode (Stratechery); now Anthropic itself has put the same figure in writing in its Series G funding announcement, adding that it has more than doubled since the start of 2026. The single-source concern is resolved. Total revenue is still a fight between reporting bases: the official announcement says $14 billion annualized, while market analyses run above $30 billion (SemiAnalysis). Treat the two as a band and you get this: one coding tool accounts for somewhere between one-tenth and one-fifth of a model company's total revenue. The foundry's house brand is not a side business; it is one of the flagship product lines.

Gauge two: labs internalizing the next vertical application. This gauge's twelve-month observation window opened only two days ago, but the supporting evidence had already happened between January and May, so the window opens at full reading. On May 7, 2026, Claude's Excel and Office integration went generally available (The Decoder); five days later Claude for Legal launched with twelve practice-area plugins and more than twenty connectors in one release, Westlaw and DocuSign among them (LawNext). On OpenAI's side, January 7 brought ChatGPT Health, which connects to medical records (Fortune), and March 24 a revamped shopping experience (CNBC). The vertical after coding didn't just arrive; it arrived in batches.

But the same gauge reads out two boundaries, and they deserve the same volume. First, the form of internalization is not a frontal assault: what Claude for Legal actually ships is plugins and connectors, and the list even includes the legal-AI company Harvey; it looks less like building a product to replace the vertical vendors and more like absorbing them onto the platform. Second, labs' first-party applications can also retreat: OpenAI announced in mid-July that it is shutting down its Atlas browser, with service reportedly ending in early August according to multiple tech outlets — this newsletter has not obtained first-hand confirmation. Labs entering the field does not mean labs winning it; their best odds remain in the layer of features that sit right next to the model application itself: call it the adjacent layer.

Gauge three: the neutral distribution layer's share. For the first time, this gauge has respectable numbers. OpenRouter, the multi-model routing service, hit 25 trillion tokens per week in May 2026, a fivefold increase in a year, and raised a $113 million Series B (Yahoo Finance). For AWS's multi-model platform Bedrock, the analyst firm SemiAnalysis estimates that more tokens flowed through it in the first quarter of 2026 than in all prior years combined, with customer spending up 170% quarter over quarter (SemiAnalysis). Absolute scale remains far below the labs' direct sales, but the growth curve is steep.

Gauge four: can third-party applications stand their ground on the labs' turf? This is the most interesting reading: the third parties are not just alive, they are growing fast. Cursor's annualized revenue went from roughly $100 million to $2 billion in about thirteen months, per market reporting (Let's Data Science), consistent with the growth trajectory this newsletter had been tracking (Stratechery); and per market reports, Replit's valuation tripled to $9 billion in half a year. These numbers directly refute the strong reading that "labs eat the adjacent layer, so third parties can't survive." So how are the third parties surviving? The answer comes two sections down: they changed species.

Why the labs can't stay out: the pure-play position is a trap in AI

With the four gauges read, answer the more fundamental question first: why don't the labs copy TSMC and simply be clean, neutral process suppliers?

Because that position, for now, does not exist in AI. Arvind Narayanan, co-author of AI Snake Oil, laid out the arithmetic in July: selling tokens is a textbook-rare case of pure perfect competition. The product is undifferentiated, switching costs are near zero, and there is no geographic protection. Capital-intensive industries that fell into this structure have ended badly: fiber optics vaporized $2 trillion; airlines have earned net margins of 2% to 4% across eighty years. And for this cycle's $4–8 trillion of AI investment to pay back at a 5% net margin (the author assumes five-year amortization) would require $1.6 to $3.2 trillion in annual revenue, roughly a quarter of world GDP (AI Snake Oil / Normal Technology). Historically, only two infrastructure industries have escaped this trap: cloud, by using lock-in to turn itself into a software business, and TSMC, by monopoly.

Note the conditions of TSMC's escape. It holds 72% of global foundry revenue today and has raised advanced-node prices four years running, on the back of process leadership plus the fact that switching foundries means a customer betting $500 million on a tape-out (Stratechery). Frontier labs have neither: open-weight models, whose parameters are published and free to download and modify, are chasing close behind, and a customer can swap models within a day. So "be the TSMC of AI" is not an available strategy; it is the road to fiber's fate. The labs have all understood this, which is why they are climbing the stack in unison: treat the API as a railway that will never make real money, and the first-party applications on top as the actual cash register. This newsletter described that structure earlier this month with Hong Kong MTR's "rail plus property" model: the train fare never pays back the track; the money is in the property built above the stations.

There is also a technical-layer reason to climb: at this stage, integration simply works better. The same model, dropped into different agent harnesses, performs wildly differently. AI researcher Nathan Lambert, who writes Interconnects, ran the same model through several harnesses and found Claude Code far ahead of Cursor's agent, and ahead of the GitHub Copilot build he was testing at the time — a qualitative test; he published no quantitative metrics (Interconnects). The quality premium of tuning model and harness together is exactly what Harvard Business School's Clayton Christensen, who originated disruption theory and the law of conservation of modularity, meant by "when the product isn't good enough, the integrator wins." This newsletter tracks the pricing-power dynamics behind this paragraph with a standing trend tree:

Infra capture era (2023-25)value eaten by chips/power/memory; labs sell tokens at a loss (2024 inference margin -94%)
Unit token economics flip (2026H1)agentic demand x cost collapse; profit in sight -- do margins hold?

Open ?

Path Adurable pricing power (Semi- Analysis): scarcity+gap -> value prices
Path Bcommoditization by default (Evans/Narayanan): identical tokens + zero switching grind margins to cost
Path Cdemand is padded (Gurley/ Chamath): subsidies + unsettled ROI, shown up when the capital window shuts

The one-line reading of this tree: per-token economics have already flipped positive, but whether the margin holds is, for now, only true at the flagship, high-difficulty tier — and building first-party applications is precisely the labs moving themselves into the tier where it holds.

The other side is climbing too: reverse integration from the application layer

The original judgment framed integration as a one-way expansion by the labs. Six months of evidence says: the traffic runs both ways.

The third parties named as lunch are training their own models in return. Cursor's self-trained Composer now carries the product as its workhorse, built on Kimi, a Chinese open-weight model, as its base (The Decoder); Windsurf shipped its own SWE-1.5 (Windsurf). Application companies reaching down into models and labs reaching up into applications are the same force expressed at both ends: nobody dares leave their lifeline in the hands of a supplier who competes with them.

That wariness is not paranoia; there is precedent. In June 2025, when reports surfaced that OpenAI was acquiring Windsurf, Anthropic cut off Windsurf's access to Claude models with less than five days' notice. Anthropic co-founder Jared Kaplan's explanation was blunt: "I think it would be odd for us to be selling Claude to OpenAI" (TechCrunch). Add Anthropic's earlier friction over blocking third-party tools from accessing Claude through subscriptions (Lex Fridman Podcast), and your-supplier-cuts-your-supply risk has now been realized more than once. The intent-level evidence goes back further still: in 2024, Sam Altman warned thin-wrapper startups that OpenAI would flatten them as it built out its own products, reported at the time under the headline "OpenAI is going to steamroll you if your startup is a wrapper on GPT-4" (Tech Startups). We made a point of checking the other direction too: as far as our verification reached, no frontier-lab chief executive has ever publicly made a commitment on the order of "we will not compete with our API customers." The sentence TSMC writes into its annual report every year has not appeared in this industry even once.

So the current equilibrium looks like this: labs go down into applications, application companies go up into models, and everyone is an integrator. The structure of this technological generation produces this; no one here is betraying anyone. Products aren't good enough yet, so integration carries a premium; selling pure process is a trap. Semiconductors took thirty years to travel from an industry where everyone integrated to specialized division of labor, and the preconditions were a mature process and standardized interfaces. At AI's stretch of that road, neither signpost has lit up.

Neutrality didn't disappear — it moved

Does the most valuable thing in the TSMC story (that pledge, "I do not compete with my customers") really have no home in AI? It has one — just not in the manufacturing layer. Look one layer up: it lives in distribution and governance, and it is already collecting money in three forms.

Form one: multi-model routing and platforms. OpenRouter's and Bedrock's growth numbers were read in the gauge section above; what they sell, stripped down, is the mirror image of the pure-play pledge: the platform builds no frontier models of its own, so when you swap models, compare prices, or mix vendors on it, its only incentive is to serve you. AWS has made this positioning an open strategy: generative AI as the new commodity building material, with customer choice as the point (Stratechery). The enterprise-side demand is structural too: Databricks CEO Ali Ghodsi has said seventy percent of his customers use more than one cloud; resistance to lock-in is itself a standing requirement (Stratechery).

The second form is the newest and the most telling: neutrality implemented as a governance structure. OpenClaw is this year's fastest-growing open-source agent tool, in wide use among agent developers; the labs raced to hire away its author, Peter Steinberger — and OpenAI got only the person. The project itself went into an independent foundation, funded by sponsorship rather than owned, with a full-time maintainer team for the first time (X / @steipete). Translated: that position was so important that the market would not allow any single lab to own it, so neutrality got institutionalized through governance. It is the cleanest pricing yet of neutrality as a scarce good.

Form three is still in its infancy: the pure-play cloud. Once cloud providers also learn to feed their own applications first, a compute supplier that does not compete with its customers acquires a unique selling point. This newsletter has recorded that "token foundry" opportunity before; it remains a hypothesis, parked on the watchlist.

One conspicuous counterexample needs a place in this picture. Google integrates everything from self-designed chips through models to its own applications, and 2026 is treating it very well: the Gemini app at 900 million monthly actives, the cloud business sprinting at roughly sixty percent annual growth, per a public compilation whose figures this newsletter has not individually verified, though the direction is consistent with multiple reports. This counterexample does not undercut the essay's judgment, because the claim here is not "integration loses" — quite the opposite: in this generation integration wins, which is why everyone integrates. And Google's integration advantage cashes out mainly inside its own ecosystem; it competes at the model and application layers but does not sell the "neutral distribution" position itself, so it is not a player in the neutrality-relocation game. This essay is about who, in an industry where everyone integrates, gets paid for the promise not to. The tug-of-war at the agent-platform layer between the vertical and neutral routes is one this newsletter also tracks with a tree:

Chat assistants (2022)language = UI
Copilot embeds (2023)in-flow, hands-on
Coding agents2024-, human sign-off
General agent platforms (2026-)cross-tool work; swappable model

Open ?

Opposing claimmeta-harness route: vendor- neutral, composable agent layer (Databricks Omnigent style)

Where we land

Two pressures, not a strategic choice, are turning frontier labs into integrated device manufacturers: this is a generational equilibrium, and blame is the wrong lens for it. Staying in the pure token-selling position is a textbook perfect-competition trap, and at the stage where the product is not yet good enough, integration carries a quality premium. The same force is pushing application-layer companies to train their own models, and the field is converging on an industry where everyone is an integrator. The genuinely valuable neutral position has not disappeared; it has moved to the distribution and governance layer: multi-model routing is already collecting revenue, and OpenClaw, this year's fastest-growing open-source agent tool, has moved into an independent, sponsor-funded foundation that institutionalizes neutrality. The pure-play cloud is still just a hypothesis on the watchlist.

>

What would prove this wrong: four things to watch. Observation windows, except where noted, are this newsletter's own twelve-month settings, and the verdict date falls when each window closes. One — if a vendor-neutral general agent layer becomes the de facto standard, with two or more large enterprises publicly adopting it as their primary execution layer, the "vertical integration wins" half falls first; this window carries over from the watch item this newsletter set in mid-July. Two — switching-cost measurement: if the next edition of Menlo Ventures' enterprise AI adoption survey (published mid-year) shows annual model-switching rates rising sharply, the quality-premium case for labs' first-party applications loosens. Three — if first-party applications' share of lab revenue turns downward, or first-party shutdowns like OpenAI's Atlas browser start appearing in the adjacent layer as well, the integrator judgment gets downgraded. Four — if a pure-play cloud produces measurable customer-defection evidence that "we do not compete with our customers" is collecting revenue at scale, the neutrality-has-moved thesis upgrades from hypothesis; if a year passes with no signal, the third form gets deleted.

What this means for you

Application-layer and agent founders. Platform risk here is documented, not hypothetical: two precedents and executives' own words are on the record. Price it like an insurance premium: build in multi-model architecture, keep the base model swappable, and lean toward use cases further from the model's adjacent layer. Note, too, that the barrier has dropped: Cursor trained its workhorse model on an open-weight base — the reverse-integration road is cheaper than it was a year ago.

Enterprise buyers. Your AI supplier is simultaneously a potential competitor to one of your product lines; in this industry that is the default, not the exception. Neutrality clauses, data isolation, and exit provisions in procurement contracts deserve to be negotiated at the specification level that semiconductor customers apply to their foundries.

Cloud and platform operators. Neutrality is a sellable asset, and it has only just started being priced. The Bedrock and OpenRouter growth curves say the "I don't build models, so you can trust me" business is growing; writing neutrality into product commitments is differentiation, not weakness.

Investors. The line "value flows to the application layer" has to be broken down by adjacency. The territory closest to the model will be attacked by the labs again and again, but attack does not guarantee success; the Atlas shutdown and Legal's connector-shaped launch show the labs' edge has boundaries. Meanwhile, the neutral distribution layer has, for the first time, revenue and valuations you can actually compare — worth a tracking slot of its own.


Written from the same research and judgments as the Traditional Chinese edition; every claim links to a primary document.