3 June 2026 · Medows · Alapan Mondal · Founder, Medows
Why Doctors Don't Trust AI (Yet)
Three reasons doctors don't trust clinical AI yet, in order of how often I hear them. None of them are about the AI being smart enough.
Alapan Mondal, B.Tech, M.Tech, IIT BBS
Founder, Medows
Three reasons doctors do not trust clinical AI yet, in order of how often I hear them.
The AI is right most of the time and wrong sometimes, and there is no good way to tell which
This is the calibration problem. A clinical AI that gives a confident answer to a hyperkalemia management question — and is correct — feels different from one that gives a confident answer about a rare differential and is wrong. The user has no signal to distinguish. Existing AI tools either provide no confidence indicator or use a meaningless one ("high confidence" applied to both cases).
The fix is not better confidence calibration on the AI side. It is better source attribution. When the AI tells me to give calcium gluconate for hyperkalemia, it should also tell me where that comes from — AIIMS protocol, ICMR critical care guidance, UpToDate-equivalent. The clinician can verify the source in seconds. The AI's confidence is irrelevant; the source's authority is what matters.
There is a deeper reason this works, beyond the verification speed. A doctor has spent years learning to weigh sources against each other — this guideline versus that one, the trial versus the local protocol, what the department actually does. Disagreeing with a source is a trained clinical act. Disagreeing with a probability is not a skill anyone has, because it is not a thing that can be done. Attribution hands the clinician a decision they already know how to make; a confidence score hands them one nobody knows how to make.
The AI doesn't know what it doesn't know
A clinician's most important skill is recognising the edge of their own knowledge. AI tools today rarely communicate this. Asked about a rare presentation, the AI confidently produces a plausible answer rather than saying "this case sits outside my training distribution, get a senior."
The fix is hard. It requires the AI to have a meaningful internal sense of its own knowledge boundaries. But the partial fix is feasible: when the AI's confidence is low (in the literal model-output-probability sense), it should escalate explicitly. The current default — confidently producing the most likely answer regardless — is the opposite of what a doctor needs.
There is a cheap diagnostic that any clinician can run on any tool they are evaluating: has it ever refused you? Ask it twenty things, three of which are genuinely at the edge — an unusual presentation, a question where the local protocol and the international guidance diverge, something outside its corpus entirely. A tool that produced twenty fluent answers has not demonstrated breadth. It has demonstrated that it does not have a "no," and a system without a "no" is not telling you anything when it says yes.
The escalation, when it comes, also has to be usable. "Consult a specialist" is not an output; it is a disclaimer wearing an output's clothes. "This is outside the protocol corpus — this is a nephrology call, and here is what they will ask you for" is a thing a resident at 3 a.m. can act on.
The AI lives outside the patient
The single most important property of a clinical tool is that it knows which patient is on the screen. A general AI assistant doesn't. The clinician has to restate the case every time, and the restatement is itself a source of error (transcription, omission, paraphrase). The AI's answer is to the restated case, not the patient. Errors compound.
Restatement is not merely lossy — it is selectively lossy, and the selection is made by the person asking. What gets typed into the box is what the doctor already considered relevant. The AI then reasons about the doctor's model of the patient rather than the patient, which means it is structurally incapable of catching the thing the doctor did not think to mention. That is the exact category of error worth catching. The same argument, worked through with a potassium of 5.8.
The fix is the workspace architecture. The AI lives inside the patient's record. The clinician doesn't restate. The AI's answer is grounded in the actual data on screen.
The fourth reason, which is not about the AI at all
Nobody lists this one first, but it is underneath the other three: the doctor carries the risk, and the tool does not.
If a resident acts on an AI suggestion and the patient is harmed, the resident is the one in front of the committee. The vendor is not. The model is not. That asymmetry is not unreasonable — the doctor made the decision — but it changes what "trust" has to mean. A clinician is not being asked to believe the tool is usually right. They are being asked to stake their registration on being able to defend, afterwards, why they followed it.
Which means the trust question is really a question about legibility. Can I explain, in a review, why this was reasonable? A sourced recommendation is defensible: I followed the AIIMS protocol, here it is. An unsourced one is not defensible at all, no matter how correct it turned out to be. This is also why the boundary should be written down rather than implied, which is why our terms say plainly that AI output is reference material and the clinical decision remains the treating physician's.
What actually earns trust is boring
These three fixes, taken together, are what makes a clinical AI tool trustable enough to use without supervision. None of them are about making the AI smarter. All of them are about making the AI's behaviour legible to the clinician.
And legibility is built out of unglamorous properties: the same question gets the same answer twice; the sources are real and open in one tap; the tool declines when it should; the numbers it fills in say where they came from. None of that appears in a demo. All of it is what a resident is actually assessing over the first two weeks, mostly without noticing they are doing it.
Trust isn't an emotional outcome. It is a property of the interface.
Author
Alapan Mondal, B.Tech, M.Tech, IIT BBS
Founder, Medows
Founder of Medows. Building doctor-side AI workspaces.
Medows is a clinical AI workspace for the doctor on rounds. Learn more or write to alapan@medows.ai / alapanx@gmail.com.