Model watch · This week (published August 27)
From the Aug 28, 2026 daily brief
Google DeepMind announced yesterday that it has completed a double-blind evaluation with outside parties (Google DeepMind, 08-27). The old dilemma was a choice of two: in a high-stakes external evaluation, either the evaluator hands over the questions and the model vendor may see the exam in advance, or the vendor hands over the weights and its intellectual property may leak. This time each side's asset was locked into a confidential computing environment on Google Cloud. The hardware first proves which code is about to run. Both parties review and approve before releasing their asset into it, and only the agreed output comes back out. The evaluator never sees the model weights, and Google never sees the questions. Cryptography's job here is proving the isolation actually held, not implementing the double blind itself. Also taking part were Singapore's official AI safety institute, the open-source community OpenMined, frontier AI audit organisation AVERI, and benchmark standards body MLCommons. AVERI's own announcement adds that the model under test was Gemini 2.5 Flash-Lite, a version number DeepMind's post omits, and that the questions came from previously unused prompts in MLCommons' safety benchmark (AVERI, 08-27). ⚠️ Four honest markers: what was announced is a method, with no evaluation results of any kind; AVERI describes it as a small-scale quantitative assessment; the two sides publish inconsistent partner lists; and "world's first" is a first-party claim.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…