Model watch · Today (published September 7)
From the Sep 8, 2026 daily brief
The independent evaluation shop Artificial Analysis reported on September 7 that OpenBMB's MiniCPM5-2B — 2.6B dense parameters, Apache 2.0 — scores 15 on its Intelligence Index v4.2, the highest of any open model below 4B and level with Qwen3.5 9B. On the banking customer-service agent test τ³-Banking it takes 21%, against 8% for the next best. On the knowledge test AA-Omniscience it scores −12, and it gets there by answering only 29% of the questions and declining the rest, at an average of 19,000 tokens of output per question (@ArtificialAnlys, 2026-09-07). (In fairness: a negative score means the test penalizes wrong answers, but we did not obtain the exact scoring rule.) ⚠️ Third-party evaluation, checkable. What is worth recording is not the score but that a small model bought it by thinking more and answering less, and converting that into latency and cost is a separate measurement.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…