Expert takes
From the Jul 19, 2026 daily brief
He went through six open-weight models' technical reports side by side and found that the training recipe for adjustable reasoning effort — letting the user decide how long and how deep the model thinks before answering — has converged on the same three-stage framework, officially graduating from research feature to factory standard (original). One concrete number: Kimi's approach cuts generation volume by about 25–30% on its own models while scores barely move on Kimi's own, unnamed internal evals (as relayed by Raschka). This supports the judgment line behind our July 16 issue on OpenAI pricing its new generation 2x higher: model capability and cost are increasingly set by how much compute goes in at inference time. "How big is the model" and "how long does it think" are now two independent procurement knobs — a small model on high effort can sometimes match a big one. Caveat: his survey spans six primary reports, but for this brief it is still a single source, and the specific numbers await checks against each original report.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
Mollick, a professor at Penn's Wharton School and the most widely read AI pragmatist among…
In his year-end interview (we only obtained the transcript last night), Google DeepMind CE…
In the same interview he decomposed the AI bubble by layer: "seed rounds for startups that…
NYU professor emeritus and prominent AI skeptic Gary Marcus argues that Chinese model capa…