Chips & semiconductors · This week (published August 31)
From the Sep 1, 2026 daily brief
Our August 31 issue covered the Oxford paper: high-bandwidth flash holds 16 times the capacity per stack of the high-bandwidth memory in use today, and dropping it straight in drags performance down, because the tail latency of flash leaves the GPU's scheduler waiting. What is new today is a second team with a different fix. Huawei, ETH Zürich and Huazhong University of Science and Technology jointly published FLINT, which likewise treats flash as a capacity tier alongside high-bandwidth memory for holding model weights. Three mechanisms carry it: a hardware burst-buffer controller that coalesces scattered reads into large sequential ones; moving flash maintenance work off the inference critical path entirely; and replacing the general-purpose address translation table that supports arbitrary writes in an SSD with a compact read-only one (Semiconductor Engineering, 08-31; the paper, arXiv 2608.25062, August 2026). Put the two papers side by side: two unconnected teams in two days, two different approaches, converging on one diagnosis of the bottleneck — inference is now limited by how much memory can hold, not by how fast the arithmetic runs. ⚠️ Both stop at the paper stage, and neither publishes a latency measurement under a real inference load. That remains the signal this line is waiting for.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
Product leaders from five electronic design automation and AI chip-design tool vendors sai…
OpenAI's enterprise announcement on September 9 recast this model from "the smartest model…
A document signed by CFO Sarah Friar gives the first numbers for Jalapeño, OpenAI's first …
Meta introduced Muse on September 8: it runs on a cloud virtual machine dedicated to each …