Research notes · Retrospective (November 2024 interview)
From the Jul 20, 2026 daily brief
Gwern, the anonymous independent researcher known for his early systematic case for the scaling hypothesis in 2020, diagnosed it this way: web corpora contain only descriptions of agent behavior, never the step-by-step decision traces — in his words, "All the agency there is is an accidental byproduct of somebody training on data" (Dwarkesh Podcast × Gwern, Nov 2024). That still has explanatory power for the 2026 reality of agents that demo brilliantly and wobble in production. It also echoes today's lead 4: the product form is evolving; the root problem of reliability is a separate thread. Time boundary: industry investment in reinforcement learning for agents has risen sharply across 2025–26; this is a diagnosis from a 2024 vantage point.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…