19 August 2026 · Medows
FDA Will Grade AI Like It Grades Doctors
The FDA wants feedback on grading generative AI medical devices the way it evaluates physicians: by competency, not just code.
The FDA asks a different question
On August 18, the FDA's Digital Health Center of Excellence published a discussion paper on how it might regulate generative AI-enabled medical devices. Not a rule. Not draft guidance. A discussion paper, open for comment until October 19, filed under docket FDA-2026-N-7874 on Regulations.gov.
The agency was careful to say this is not policy yet. Acting Commissioner Kyle Diamantas framed it as a race for standards: "The United States must lead in shaping how this technology is developed and used safely and responsibly." Center Director Michelle Tarver called it "a transparent process to inform the development of an approach that safeguards patients."
That's the usual language of a regulatory rollout. What's not usual is the model the FDA is reaching for.
Graded like a resident, not like a device
Traditional software as a medical device gets evaluated once, then locked. Change the code, resubmit. Generative AI doesn't sit still. A large language model behind a diagnostic tool or a documentation assistant can behave differently depending on the prompt, the patient population, the week.
So the paper proposes something closer to how hospitals already handle uncertainty in a clinician: competency assessment. In the FDA's own words, premarket evaluation would be "built on the concept of competency assessment, inspired at a high level by how physicians are trained and evaluated." Alongside that, the paper lays out a two-axis framework for risk, a mix of non-clinical benchmarking and clinical confirmation, risk-proportionate postmarket monitoring, and separate considerations for foundation models and agentic AI systems that act rather than just answer.
Read that plainly: the FDA is proposing to treat an algorithm less like a static instrument and more like a trainee whose competence has to be demonstrated, checked, and rechecked. DHCoE Director Rick Abramson called the paper an attempt to "advance the frontiers of regulatory science."
Why this matters on the ward, not just in Washington
Doctors have been living with a version of this problem for two years. An ambient scribe drafts a note. A diagnostic assistant flags a finding. A summarization tool condenses a chart. None of these tools come with a transcript of exactly how they earned the doctor's trust, and none of them get re-examined the way a resident does on rotation.
If the FDA's competency framework holds, that changes. A generative AI device inside clinical workflow would need to show its work: what it was tested against, where it holds up, where it doesn't, and how that gets rechecked after deployment, not just before.
That is a higher bar than most AI tools clinicians use today clear. Point tools that summarize or transcribe without exposing their failure modes will have a harder time under a framework built around demonstrated, checked competence. Tools built around verifiable output, sources a doctor can actually trace back, will have an easier one.
None of this is settled. It's a discussion paper, and the comment period runs through mid-October. But the direction is worth noting: the agency writing the rules for AI in medicine is starting to ask the same question a chief resident asks of every new doctor on the ward. Not "does it work," but "can you show me it works, and will you still be right next week."
Sources
Medows is a clinical AI workspace for the doctor on rounds. Learn more or write to alapan@medows.ai / alapanx@gmail.com.