Model watch
From the Aug 24, 2026 daily brief
one, measure the cache hit rate before you look at throughput — once the hit rate collapses, every throughput number after it is wasted arithmetic. Two, if you want main-memory offload, main memory has to be substantially larger than the GPU's cache capacity, because the design copies on write: every entry written into GPU memory is written into main memory at the same time. Too small and you have effectively not enabled it. The measurer suggests 1.5 to 3 times, though the original does not explain how either end of that range was derived. ⚠️ Both come from today's single-source measurement, with no third-party re-run. Separately, the academic papers we swept this week held nothing new worth its own item, so this column has nothing new for you this issue — we do not dress an evergreen concept up as news.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…