14 August 2026 · Medows
Clinical AI's Real Danger: What It Skips
The largest clinical AI safety study yet found most severe errors are omissions, not wrong answers, and purpose-built tools beat general chatbots.
It's 2 a.m. and a resident is three patients deep into a differential nobody has time to double check. They open a chatbot, type the case, and take the first answer that sounds right. That moment, repeated across thousands of wards every night, is exactly what a new study tried to measure.
The largest test of its kind
Last week, researchers from Stanford Medicine, Harvard Medical School, and the ARISE AI research network published NOHARM, a 1,100-case benchmark built from real primary-care-to-specialist consultation scenarios. More than 50 researchers and 29 board-certified physicians built it, producing 12,747 expert annotations across 10 medical specialties. They ran 20 general-purpose large language models and 4 clinical AI tools through it and scored every recommendation for harm potential, not just correctness.
The headline number is not the scary one
The study found that applying AI-generated recommendations directly, without a clinician's review, carried potential for severe harm in up to 24.6 percent of cases. That number will get quoted everywhere. The more useful number is buried one line down: more than 80 percent of the severe errors were errors of omission, not errors of fact. The AI wasn't confidently wrong. It was confidently incomplete, leaving out a step, a red flag, a differential that mattered, while sounding just as certain as it did on the parts it got right.
That distinction matters for how a doctor should read any AI output. A wrong answer is something you can argue with. A missing answer is invisible until the patient you didn't test for shows up two days later.
Purpose-built tools did better, and so did paying attention
Two other findings stood out. Clinical AI tools built specifically for medical consultation outperformed general-purpose frontier models by a wide margin, which is not surprising on its own, but is a useful data point against the assumption that any capable chatbot is "good enough" for clinical use. And in the physician-AI teaming arm of the study, doctors who had AI assistance available frequently ignored good suggestions the AI offered, and as a result underperformed several AI systems running unsupervised. The tool was often right. The doctor, skimming under time pressure, didn't take the suggestion seriously enough to check it.
The marketing raced ahead of the finding
Within a day of NOHARM's release, at least two vendors, Doximity and AMBOSS, put out press releases claiming a first-place finish on the same study. Both are technically defensible: the benchmark scores multiple sub-samples and metrics, so different cuts produce different leaderboards. Neither claim is false. Both illustrate the actual problem the study surfaced: safety claims move faster than doctors can verify them, and a ranking in a press release is not the same as a citation a physician can check at 2 a.m.
Why this matters on the ward
None of this is a verdict on any single product. It's a reminder of where clinical AI actually fails right now: not in getting facts wrong, but in quietly leaving things out, and in being trusted, or ignored, at the wrong moments. A workspace that surfaces its sources and reasoning at the point of care, rather than a chat window that hands back a paragraph, gives a doctor something to check against, which is the only real defense against an omission you'd otherwise never notice. That is the gap NOHARM actually measured, and it is worth more attention than the leaderboard.
Sources
- First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations (arXiv)
- Doximity Outranks OpenEvidence, Frontier Models in Independent Stanford-Harvard Study of Clinical AI Safety (Doximity press release)
- Ranked #1 in Stanford-Harvard NOHARM Study: AMBOSS AI Mode (AMBOSS Newsroom)
Medows is a clinical AI workspace for the doctor on rounds. Learn more or write to alapan@medows.ai / alapanx@gmail.com.