Papers and technical work — only the ones that move an industry judgment.
69 pieces
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…
Six academic papers reached the reading list last night and none was finished today. Three…
The independent evaluation shop Artificial Analysis reported on September 7 that OpenBMB's…
Per Import AI's account of a Google DeepMind paper: 100 Gemini 3.1 Pro agents worked 71 pr…
Of the 411 academic papers that reached last night's reading list, not one was finished to…
Epoch AI, an independent research organization that tracks AI compute and capability trend…
292 academic papers went onto last night's reading list and not one was finished today. We…
On August 14 we recorded a post: an evaluator with early access to Z.ai's open-weight flag…
On December 20, 2025, OpenAI published "Monitoring Monitorability", which turned how well …
CyberGym is a public benchmark from a Berkeley team, with 1,507 real vulnerabilities acros…
First the background: most large models today use a mixture of experts design. Rather than…
The Allen Institute for AI published BenchMIRT on September 1. It uses a multidimensional …
Not for want of material — one item today would have qualified: Google Research's next-gen…
What this simulator measures is how a model behaves on hardware, not how the model itself …
Our August 28 issue covered GLM-5.3-Flash's pricing and its efficiency split; what is new …
One user reported that Qwen3.8-Flash broke down on multi-turn tracking at low precision, a…
Four researchers at Northeastern University, a research university in Boston, posted a pap…
one more question when you buy inference capacity — who shares memory with the GPUs you ar…
Google DeepMind announced yesterday that it has completed a double-blind evaluation with o…
anyone writing evaluation requirements now has a workable practice to cite. AVERI's announ…
Nobody read any of the 136 academic papers from last night's sweep in the original today. …
A methodological turn in the race to build AI agents that write GPU kernels: the compiler …
The standard way to evaluate code-rewriting ability checks only whether behaviour is corre…
The same measurement holds one section you can act on directly. DP-attention splits the at…
one, measure the cache hit rate before you look at throughput — once the hit rate collapse…
Together AI serves inference on open-weight models, and it published two comparisons on th…
— "a cheap open-weight model plus orchestration beats expensive closed-source" is its comm…
Our records hold a September 2025 safety study from OpenAI and Apollo Research (an indepen…
The first is coming back empty-handed. We ran a targeted search for independent replicatio…
Start with who this is. Z.ai (Zhipu) is a Chinese frontier model team, developer of the op…
we have tracked one judgment for a long time — that the real constraint on AI capability i…
Not one of the 106 arXiv papers and 50 paper abstracts that arrived overnight has been jud…
A team reported on the blog of the model hosting platform Hugging Face that their constrai…
An ordinary large language model generates token by token, which is why its chain of thoug…
In March 2025, researchers at OpenAI published a finding: letting a weaker model read a st…
Stella Biderman is executive director of EleutherAI, the open-source AI research organisat…
Distillation means using one model's output (the teacher) to train another (the student). …
On August 10, a Google team posted AMIE (Video) to arXiv — a Gemini-based multi-agent syst…
The researcher Ziqian Zhong started from a mechanical detail: Claude Code puts the user's …
"Third axis" is the paper's own framing; we do not endorse it. "Explorative Modeling: Unlo…
The eval house Artificial Analysis runs its AA-Omniscience board on an unusual design: wro…
Teortaxes ran a two-path check on DeepSeek's official disclosures for its late-2024 V3 mod…
First, which yardstick produced "one point apart": the composite intelligence index from t…
Two lines from the K3 launch materials (verbatim, as relayed by Teortaxes): late in develo…
Overnight's 15 new papers are still unprocessed and this column has no single-day incremen…
A newer route in AI training is rubric-based rewards: instead of training a separate rewar…
Epoch AI, the AI data-research organization, and the evaluator METR released MirrorCode, a…
The current mainstay of post-training (the stage after base training where a model is taug…
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…
The Jacobian conjecture, posed in 1939, is a famous problem in algebraic geometry. It says…
Our July 21 Research Notes covered this empirical study (gains from optimizing an agent pi…
Alex Zhang, author of the RLM framework paper, argues that carefully designed task orchest…
Raia Hadsell, VP of Research at Google DeepMind, argued in a talk that the recipe that bui…
An empirical study of 25,264 pull requests (code-change proposals) opened by AI agents acr…
Three mainstream "don't touch the model, optimize the agent workflow" methods went through…
François Chollet, creator of the Keras framework, once gave the claim "large models memori…
Gwern, the anonymous independent researcher known for his early systematic case for the sc…
The conclusion first: evaluating AI with AI-generated data has a structural blind spot — t…
For problems with no answer key, sample many solutions from the model, take the majority-c…
Yesterday's July 18 issue introduced MemCon from a UCLA-affiliated team in these research …
A Meta FAIR-affiliated team's March paper, Principia, argues that current math benchmarks …
Oracle, in a 13-author technical report on July 14, builds agent memory as a database-nati…
A paper from the team that includes Turing Award winner Yoshua Bengio shows that an AI age…
Practitioners reading Thinking Machines' official model card for Inkling flagged a rare se…
Databricks benchmarked models on engineering tasks against its own multi-million-line code…
The honest note on the academic side: this week's academic main course is the five peer-re…
Just an email address, unsubscribe anytime. This is the only thing we ask of you.