SecondSourceJudgment rebuilt from primary sources
Product & chips · Sep 1, 2026

A second paper in two days on putting models into flash memory, from a completely unrelated team.

Chips & semiconductors · This week (published August 31)

From the Sep 1, 2026 daily brief

Our August 31 issue covered the Oxford paper: high-bandwidth flash holds 16 times the capacity per stack of the high-bandwidth memory in use today, and dropping it straight in drags performance down, because the tail latency of flash leaves the GPU's scheduler waiting. What is new today is a second team with a different fix. Huawei, ETH Zürich and Huazhong University of Science and Technology jointly published FLINT, which likewise treats flash as a capacity tier alongside high-bandwidth memory for holding model weights. Three mechanisms carry it: a hardware burst-buffer controller that coalesces scattered reads into large sequential ones; moving flash maintenance work off the inference critical path entirely; and replacing the general-purpose address translation table that supports arbitrary writes in an SSD with a compact read-only one (Semiconductor Engineering, 08-31; the paper, arXiv 2608.25062, August 2026). Put the two papers side by side: two unconnected teams in two days, two different approaches, converging on one diagnosis of the bottleneck — inference is now limited by how much memory can hold, not by how fast the arithmetic runs. ⚠️ Both stop at the paper stage, and neither publishes a latency measurement under a real inference load. That remains the signal this line is waiting for.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section