21 August 2026 · Medows
Diagnostic AI's 95% Accuracy Problem
New research: data leakage inflates AI schizophrenia-detection accuracy by up to 30%, exposing why near-perfect diagnostic AI claims deserve real scrutiny.
A number that looks too good
Papers claiming AI can detect schizophrenia from a brain scan almost always report accuracy north of 95 percent. That is the kind of number that ends up in pitch decks and news headlines. A new methodological review says most of it does not hold up.
What the review found
Frigyes Sámuel Rácz and Gábor Csukly published their analysis in Translational Psychiatry on August 15, 2026, after an earlier version circulated as a medRxiv preprint. They went back through the published literature on EEG based schizophrenia detection and checked how each study split its data between training and testing.
The finding: 65 percent of the papers they reviewed had a pipeline error serious enough to inflate the reported accuracy. More than half split their data by "epoch," a short slice of a single patient's recording, rather than by patient. That sounds like a technical footnote, but it changes what the model actually learns. If an algorithm trains on one segment of Patient A's EEG and gets tested on a different segment from the same Patient A, it can hit a near perfect score by recognizing that patient's individual signal, not the biological pattern of schizophrenia.
When the researchers corrected for that leakage, reported accuracy fell by as much as 30 percentage points. A headline result of 95 percent can, in reality, be closer to 65.
Why a doctor on the ward should care
Nobody is deploying an EEG schizophrenia classifier on rounds this week. What makes this relevant to any doctor using AI tools is the pattern, not the specific test. A diagnostic AI accuracy number is a claim about one dataset and one way of splitting it. It is not automatically a claim about how the tool will perform on the next patient walking through the door. The paper is a reminder that "95 percent accurate" is marketing language until someone shows the split behind it.
This is exactly the gap that shows up at the bedside. A tool that scores well in a validation paper and then behaves inconsistently in real use is not a contradiction. It is often the leakage problem surfacing downstream, after the purchasing or deployment decision has already been made.
The Medows angle
This is why Medows treats every AI output as something to check, not something to trust on the strength of a benchmark. A workspace built for the doctor's actual shift should show its sources and reasoning inline, so a clinician can look at the claim and the evidence in the same view instead of taking a vendor's accuracy slide at face value. Verifiable beats impressive. A tool that says "here is what I found and where" is more useful on a ward than one that says "trust the 95 percent."
The uncomfortable part of this study is not that one narrow EEG test was overstated. It is that the same data-splitting mistake likely sits underneath other clinical AI claims nobody has re-checked yet. Any near-perfect number in a diagnostic AI paper is worth a second look before it changes how a patient gets triaged.
Sources
Medows is a clinical AI workspace for the doctor on rounds. Learn more or write to alapan@medows.ai / alapanx@gmail.com.