4 August 2026 · Medows
Which AI Can a Doctor Actually Trust?
A new leaderboard scores 18 AI models on medical safety. The general chatbots doctors reach for aren't the safest ones in the room.
A scorecard nobody asked for, until now
On July 27, 2026, a research network called ARISE (clinicians, researchers, and builders spanning several academic medical centers) launched something clinical AI has needed for a while: a public, living leaderboard called MAST, the Medical AI Superintelligence Test. It scores AI models on the things that actually matter on a ward: diagnosis, management, safety, radiology, multimodal reasoning, and agentic workflows. By August 4, the leaderboard covered 18 models.
The question MAST asks on its homepage is blunt: "Which AI can you trust for medical questions?"
The answer, so far, is not the one you'd guess.
The general chatbots are not the safest ones in the room
On MAST's overall composite ranking, the familiar names sit at the top: GPT-5.5 leads at 61.6%, Gemini 3.5 Flash follows at 59.4%, and Claude Opus 4.7 at 59.0%.
But composite scores flatten a distinction that matters more at the bedside: safety. MAST includes a dedicated safety benchmark, First Do NOHARM v2, built on the earlier NOHARM methodology, real physician-to-specialist consultation cases graded by specialists for whether an AI's recommendation could hurt a patient. On that narrower, sharper measure, the ranking reshuffles. Purpose-built clinical tools such as LiSA 2.5 (86.2%), Doximity Ask 6.1 (84.5%), OpenEvidence (80.0%), and Glass 5.6 Max (79.7%) outscore every general-purpose frontier model on safety. Claude Opus 5 comes closest among the generalists at 73.7%, GPT-5.6 Sol scores 70.1%, GPT-5.5 scores 70.0%, and Gemini's entries trail at 62.6% (3.1 Pro) and 61.9% (2.5 Pro).
MAST's own framing of the results is that "there is no single model that dominates every clinically relevant capability." A model can reason well and still be worse, specifically, at not recommending something dangerous.
That distinction was the whole point of the earlier NOHARM benchmark this leaderboard builds on: a study of 100 real physician-to-specialist cases across 10 specialties found some general LLMs produced severely harmful recommendations in more than 20% of cases, and that over 75% of the severe errors were omissions, not wrong answers, but the thing nobody said. The models sounded confident. They just left something out.
Why this is the right story to watch, not just the right leaderboard
None of this means the tools at the bottom are bad, or the tools at the top are done. MAST itself says it's in preview and scores will move as validation continues. What it does mean is that "smart" and "safe to hand a patient's chart to" are turning out to be different axes, measured differently, and that the gap between them is now something you can actually look up instead of guess at.
That is the argument for verifiable AI in a clinical setting, made in someone else's data. A tool that summarizes a consult well but omits a contraindication is not a productivity win, it is a liability wearing a productivity win's clothes. The doctor on rounds does not have time to independently re-verify every AI-generated line against the chart. The tool has to be built, from the ground up, so that what it says is checkable against what is actually in the patient's record, not just plausible-sounding.
Leaderboards like MAST won't settle which model belongs in a ward workflow. But they are the first sign that the industry is starting to measure the thing that actually matters to a doctor at 2 a.m.: not "does this sound right," but "is this safe, and can I check it."
Sources
- MAST: Medical AI Superintelligence Test (leaderboard)
- MAST technical results and safety scores
- Introducing MAST v1.0: A Framework for Evaluating Medical AI Across Clinical Capabilities
- The Medical AI Superintelligence Test and NOHARM: A New Framework for Assessing Clinical Safety in AI Systems, Cytel
Medows is a clinical AI workspace for the doctor on rounds. Learn more or write to alapan@medows.ai / alapanx@gmail.com.