SecondSourceJudgment rebuilt from primary sources
Research · Aug 12, 2026

AI video consultations get their first randomised trial: clinical scores on par or better, and patient preference splits honestly down the middle.

Model watch · Today (research, event August 10)

From the Aug 12, 2026 daily brief

On August 10, a Google team posted AMIE (Video) to arXiv — a Gemini-based multi-agent system. An OSCE, or Objective Structured Clinical Examination, is the standard exam design used in medical education: trained actors play standardised patients while candidates rotate through stations and are scored. Here the system was compared against family physicians on live video consultations under a randomised OSCE: 30 primary care physicians, 15 patient actors, 100 clinical scenarios, three arms (AI video, AI text-only, physician video). The results come in three tiers of strength. Clinical evaluators rated AMIE (Video) "on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination." Patient preference split: the actors preferred the AI's way of assessing and explaining conditions, while physicians were preferred for rapport and partnership building. And the authors list three shortfalls themselves — "limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements" (arXiv 2608.09861). Eric Topol, the cardiologist who directs the Scripps Research Translational Institute, reposted it the next day and cooled it down unprompted: "Of course this will require assessment in real world medicine," while pointing out that video adds back part of the physical examination — a dimension a text-only consultation structurally cannot reach (the post). ⚠️ This is a preprint, not peer-reviewed, not reproduced, and the same party set the exam and sat it: the rubric and taxonomy were defined by the Google team. Against that, the clinical scoring was done by clinical evaluators and the authors volunteered three failure modes, which offsets part of the concern; and standardised patients are not real patients. Our suggestion is to swap the question out: do not ask whether AI will replace doctors, ask how you design a product around that split in preference — explanation to the AI, relationship to the human. That is a concrete division-of-labour hypothesis, and it is something you can take into product design or put to a vendor.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section