Chips & semiconductors · This quarter (event August 20)
From the Sep 12, 2026 daily brief
An inference chip is a custom part that handles only the stage after training, when a model actually answers questions, giving up generality for efficiency. On August 20 an account posted a three-part challenge to Etched, a startup in that lane: go the high-bandwidth-memory route and you cannot beat GPUs and TPUs; go the on-chip static memory route and you cannot beat Groq and Cerebras; and, in his words, "Low-voltage inference is BS: you need energy to move data; this is physics" (Bing Xu on X, 2026-08-20). ⚠️ We have not verified who is behind this account. This is our first time recording it. Do not treat it as an insider because it sounds like one. What carries the weight is two pieces of material we already hold, not him. The third sentence does not hold: moving data does cost energy, but that is a floor, not a constant, and pushing the floor down is exactly what the industry is doing. d-Matrix, another inference chipmaker, says its 3D-stacked memory test chip reaches roughly 0.4 picojoules per bit in the worst case against 3 to 4 picojoules for high-bandwidth memory, an order of magnitude apart (d-Matrix blog, 2026-03-16; ⚠️ a vendor's first-party claim, not independently measured). The second sentence is actually strengthened, in a direction he may not have counted on: at its March 2026 conference Nvidia unveiled a next-generation inference chip integrating Groq's technology, with 500 MB of static memory per chip and 1.2 petaflops at FP8 (More Than Moore, 2026-03-16). ⇒ The incumbent on the static-memory route is no longer just Groq, it is Nvidia, and a startup on that route is in a worse spot than he described.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The blueprint sets out six principles covering AI literacy, age-appropriate protections, p…
If this inference holds, part of Anthropic's compute expansion is propped up by a chip sup…
This week's most important chip-side readings are CoreWeave's self-reported GPU prices and…
It connects GPT-6 Astra to a US legal research index covering more than 230M URLs, plus le…