Models & research · Named benchmark, 2026-07
From the Jul 17, 2026 daily brief
Databricks benchmarked models on engineering tasks against its own multi-million-line codebase: Sonnet 5 is about 1.7× cheaper than Opus 4.8 per token, yet costs more per completed task — $2.09 per task for Sonnet 5 versus $1.94 for Opus 4.8 — because the cheaper model retries more rounds and completes 6 percentage points fewer tasks (Databricks blog, 2026-07; The Register also covered it independently on 07-13). The one-line takeaway for buyers: compare cost per completed task, not the sticker price.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…