Model watch
From the Sep 22, 2026 daily brief
We haven't finished the 164 papers we pulled in overnight, and the academic abstracts held no same-day paper news; we don't dress up evergreen concepts as news. The item below fills the column from the past month and carries its original date.
1. [This month] (posted September 11) Same model, different harness: scores on a science-task benchmark went from 3 to 11. On another leaderboard, first place rests on 4 of 1,270 questions. The harness is the layer of code around a model that actually reads files and runs commands; our September 21 issue introduced it. The anonymous Chinese AI watcher @teortaxesTex posted that on Terminal-Bench-Science, a benchmark for agents completing science tasks in a terminal, DeepSeek V4.1 solved 3 of 70 tasks with the Terminus 2 harness and 11 with OpenAI's Codex harness (@teortaxesTex, 2026-09-11). He also pointed out that V4.1 ranks first on the agentic-coding subcategory of the public leaderboard LiveBench, and that the ranking rests on 4 Python questions out of 1,270 (@teortaxesTex, 2026-09-11). ⚠️ Both come from the same anonymous account on the same day, the 11-task run has no screenshot, and nothing was controlled, so treat the jump from 3 to 11 as an upper bound. He's a DeepSeek fan who debunked a ranking that flatters DeepSeek, which eases the worry about motivated bias. ⇒ If you choose models based on leaderboard rankings, read the rank together with the harness and the number of questions in the subcategory; when neither is disclosed, the cheapest check is to rerun it in your own harness.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
Our overnight sweep covered 47 authors and turned up nothing new; we don't dress up evergr…
The experiment takes representative open-source model designs and training recipes from ea…
We haven't read any of last night's 145 new papers or 49 paper summaries, so there's no ne…
Nothing new landed in this section, and we would rather leave the column empty than pass o…