SecondSourceJudgment rebuilt from primary sources
Product & chips · Aug 31, 2026

Sixteen times the capacity per stack in flash memory — and dropping it straight in starves the GPU. Oxford proposes the compromise.

Chips & semiconductors · This week (published August 30)

From the Aug 31, 2026 daily brief

A research team at the University of Oxford published a paper: at comparable bandwidth, high-bandwidth flash memory holds 16 times the capacity per stack of the high-bandwidth memory in use today. That looks like the answer to inference workloads that cannot fit a large model in memory. Swapping it straight in drags performance down badly, because the tail latency of flash leaves the GPU's scheduler waiting. Their approach mixes the two kinds of memory and uses prediction to keep the slow half off the critical path (Semiconductor Engineering, 08-30; the paper, IEEE Computer Architecture Letters, August 2026). Why it is worth keeping: items 1 and 2 of the main line are both about the money in inference, and the physical floor under that cost is whether the model fits in memory. Capacity at a sixteenth of the cost is a big number, and what you pay for it is latency. Nobody has finished that calculation, and until someone does, every argument that memory is about to get cheaper is missing half its case. The signal that it has been finished is a published latency measurement of the mixed architecture under a real inference load, rather than one more capacity multiple.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section