7 September 2026 · Medows
OpenEvidence Hits 100% on MedQA. So What?
OpenEvidence's new AI models score a perfect 100% on MedQA. That solves a test, not the doctor's real problem on rounds.
A resident on the 6 a.m. round has thirty seconds between patients. She does not need a five-minute literature review. She needs one number, right now, with a source she can point to if someone asks. That gap, speed versus rigor, is exactly what OpenEvidence just tried to solve.
What launched
On September 3, OpenEvidence introduced a family of four medical AI models, each named after a figure in evidence-based medicine. Osler is the fastest, answering in about five seconds. Sackett takes about thirty seconds and weighs the evidence more deeply. Snow is the slowest of the three public models, running a full literature investigation over roughly five minutes before producing a report. All three are free to verified clinicians today, on the web and the OpenEvidence apps.
The fourth, Darwin, is not public. It is in research preview, available only to institutional partners and researchers who apply for access. OpenEvidence says that is because of dual-use risk in virology, immunology, and genetics research, not because the model itself is unready.
Founder and CEO Daniel Nadler put the design logic simply: "Every model in the family is held to the same standard of clinical accuracy. What varies is time: how long a model thinks, and how deep it searches."
The number everyone will repeat
Darwin scored 100% on MedQA, a 660-question benchmark, the first model to do that. It also posted 72.8% on MedXpertQA, 82.7% on HealthBench Professional, and 0.872 on NOHARM, a severity-weighted safety score. Those are real, and they beat the prior best models by wide margins on the harder two.
But MedQA is built from board-exam-style questions: one patient, one clean vignette, one correct answer among five choices. A perfect score there proves a model can pass the test doctors already passed years ago. It says less about the actual shape of a ward shift, where the history is incomplete, three consults disagree, and the "one correct answer" doesn't exist until you've chased down the potassium from two hours ago.
Naming the right ancestors, solving half the problem
The naming is not an accident, and it is the part of this launch worth taking seriously. OpenEvidence describes Osler as the figure who moved medicine from the lecture hall to the bedside, Sackett as the founder of evidence-based medicine, and Snow as the founder of modern epidemiology. OpenEvidence is explicitly claiming that lineage: fast, sourced, evidence-graded answers instead of a confident guess.
That is the right instinct. A doctor asking "what's the evidence on X" deserves a model that shows its work and grades the strength of what it found, not just a fluent paragraph.
Where it stops short is scope. Osler, Sackett, and Snow are still answer engines: you ask, one model tier picks up the question, it answers. Each query is its own island. A shift is not a query. It's a continuous, changing case: labs that land mid-round, a handover note from last night, an order set you're about to write, a family asking a question in the hallway. A doctor who gets a fast, well-sourced answer to one question still has to carry that answer, unaided, into the next six things happening on the ward, and the tool has no memory of any of it once the answer is delivered.
Why this matters beyond one product
This is the same gap we built Medows around: verifiable answers are necessary but not sufficient. A model that cites its sources and grades its evidence is doing something real. But bolting that model onto a workflow made of a dozen disconnected screens doesn't remove the disconnection, it just makes each individual screen answer faster. The doctor still has to be the integration layer, stitching a fast answer from one tool into a note in another, a plan in a third, and a handover in a fourth.
Darwin's perfect score is a genuine technical result. It is not evidence that the bedside problem is solved. The bedside problem was never "can a model find the right answer fast enough." It's "can the right answer show up already connected to this patient, this shift, this round," without the doctor doing the stitching by hand. That's a workspace problem, not a benchmark problem, and no single query, however fast or however well-sourced, closes it on its own.
Sources
Medows is a clinical AI workspace for the doctor on rounds. Learn more or write to alapan@medows.ai / alapanx@gmail.com.