14 September 2026 · Medows
Autonomous AI Wins Benchmarks, Not Wards
Dueling STAT essays debate if autonomous AI beats doctors. The benchmarks are real, but accountability on the ward isn't in the data.
Two doctors published dueling essays in STAT News on the same day, September 9, 2026. One said autonomous AI is already better than doctors, even doctors using AI. The other, the CEO of the American Medical Association, said AI should never run the conversation on its own. Neither is wrong about the evidence. They disagree about what the evidence is for.
The case for the machine
Ezekiel Emanuel and Abe Baker-Butler, both at the University of Pennsylvania, laid out the numbers plainly. ChatGPT beat physicians at differential diagnosis by 18 percentage points (92% versus 74%). Microsoft's AI Diagnostic Orchestrator reached the correct final diagnosis 4.02 times more often than physicians did on hard case reports, 80.4% versus 20%. MIRA, a prescribing model, chose guideline-concordant treatment 35 percentage points more often than doctors. In a Stanford study, an autonomous AI system reached a stable insulin dose for diabetic patients in 15 days. Doctors managing the same task hadn't gotten there after eight weeks.
The authors go further than "AI helps doctors." They count 13 studies since January 2024 comparing autonomous AI against physicians working with AI assistance. Nine show the autonomous system still wins, by an average of 21.3 points on diagnostic accuracy in the cases they cite. They also point to empathy scores: 13 of 15 studies found patients rated AI communication as more empathetic than a human clinician's, and in trials with Google's AMIE, patient actors reported feeling more at ease (97% versus 65%) and more listened to (95% versus 72%) than with primary care physicians.
Their conclusion is that a doctor "supervising" an AI that already outperforms them is a fig leaf, not a safeguard. If the AI is right more often alone, they argue, the honest move is to let it run the five core tasks: history-taking, differential diagnosis, test selection, treatment choice, chronic disease management, and build proper licensure and liability rules around that.
The case for the human
John Whyte's response, published the same day, doesn't dispute the benchmark numbers. It disputes what they measure. AI can interpret images, summarize a chart, suggest a diagnosis, he writes, but that isn't the same as practicing medicine. His line: "AI can inform those conversations. It may make them better. But it should not conduct them on its own." He points to the more mundane, less debatable win: AI cutting time doctors spend searching for information, documenting encounters, and doing administrative work, freeing them for the parts of the job that are actually about judgment and responsibility.
What the benchmark doesn't carry
Both essays are right about something, and the gap between them is the real story. A model that hits 92% on a diagnostic vignette was tested on a case that already had an answer. A ward doesn't. Someone has to decide what to test for next when the picture is still forming, own the call when the guideline doesn't fit the patient in front of them, and be named when a family asks what happened. None of the studies cited measure that, because a benchmark can't contain it.
That's the actual argument for keeping a doctor in the loop, not sentiment about bedside manner. It's not that physicians are more accurate than the software at pattern-matching a fixed case; on the evidence assembled here, often they aren't. It's that accountability doesn't transfer to a model, and a hospital that hands over the five core tasks still needs one human who can be asked "why" and answer for it at 3 a.m.
That's the design question underneath both essays: not whether AI is capable, everyone now agrees it is, but whether it sits inside the doctor's workflow, checkable and answerable to them, or sits above it. The benchmark race will keep producing better numbers. The ward will still need someone to sign the chart.
Sources
- AI alone can be even better at medicine than AI-assisted doctors, Ezekiel J. Emanuel and Abe Baker-Butler, STAT News, September 9, 2026
- AMA CEO: AI won't replace doctors, it will work alongside them, John Whyte, STAT News, September 9, 2026
Medows is a clinical AI workspace for the doctor on rounds. Learn more or write to alapan@medows.ai / alapanx@gmail.com.