Research Notes · This week Trend (published around 07-20)
From the Jul 22, 2026 daily brief
Alex Zhang, author of the RLM framework paper, argues that carefully designed task orchestration (the loops and tool-dispatch logic around the model) can reduce superficially different tasks to similar execution traces, letting training on short tasks generalize to tasks 8 to 32 times longer (X / @a1zhang, 07-20). Commentator swyx supplied the dark side: you don't need to train on the test itself — train on data that merely looks like the test and you can farm impressive scores. The same mechanism is at once a capability source and a benchmark-contamination source (X / @swyx, 07-21). Caveat: the 8–32× figure is from a single paper, not independently reproduced. The takeaway question: when an agent benchmark score jumps, first ask whether the jump is in the model or in the orchestration.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…