SecondSourceJudgment rebuilt from primary sources
Product & chips · Sep 9, 2026

OpenAI says its first in-house inference chip delivers 1.5 to 1.9 times the throughput per watt of "the commercial systems tested" — and does not say which systems those were.

Chips & semiconductors · Today (published September 8)

From the Sep 9, 2026 daily brief

A document signed by CFO Sarah Friar gives the first numbers for Jalapeño, OpenAI's first custom inference chip. In the company's own words:

Jalapeño, our first custom inference chip, extends that work into hardware. In InferenceX tests across three public models, it delivered 1.5 to 1.9 times as much peak token throughput per watt as the commercial systems tested, using rated chip power to normalize the comparison. End-to-end latency was 1.7 to 3.6 times lower. We plan to begin deploying it by year-end alongside accelerators from NVIDIA, AMD and other partners.

(OpenAI, 2026-09-08). Why per watt rather than per chip: in an industry where power is already the binding constraint, how many tokens a watt yields is closer to a real cost measure than how fast one chip runs. ⚠️ Three discounts stack. The commercial systems are unnamed, so the multiple has no comparable object attached to it. The normalization uses rated chip power rather than measured draw, which favors whichever side actually consumes less than its rating — the document says so itself. And every figure is self-reported with no third-party reconciliation. The text says deployment begins by year-end alongside Nvidia's and AMD's accelerators, and says nothing about whether the chip will be sold externally. Buying in and building your own are the challenger's two routes, and today each produced one reading: one route pays an entrance fee in equity, the other claims a performance advantage that only one party is asserting.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section