Chips & semiconductors · This week (published August 30)
From the Aug 31, 2026 daily brief
A research team at the University of Oxford published a paper: at comparable bandwidth, high-bandwidth flash memory holds 16 times the capacity per stack of the high-bandwidth memory in use today. That looks like the answer to inference workloads that cannot fit a large model in memory. Swapping it straight in drags performance down badly, because the tail latency of flash leaves the GPU's scheduler waiting. Their approach mixes the two kinds of memory and uses prediction to keep the slow half off the critical path (Semiconductor Engineering, 08-30; the paper, IEEE Computer Architecture Letters, August 2026). Why it is worth keeping: items 1 and 2 of the main line are both about the money in inference, and the physical floor under that cost is whether the model fits in memory. Capacity at a sixteenth of the cost is a big number, and what you pay for it is latency. Nobody has finished that calculation, and until someone does, every argument that memory is about to get cheaper is missing half its case. The signal that it has been finished is a published latency measurement of the mixed architecture under a real inference load, rather than one more capacity multiple.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
Product leaders from five electronic design automation and AI chip-design tool vendors sai…
OpenAI's enterprise announcement on September 9 recast this model from "the smartest model…
A document signed by CFO Sarah Friar gives the first numbers for Jalapeño, OpenAI's first …
Meta introduced Muse on September 8: it runs on a cloud virtual machine dedicated to each …