Policypublished

FDA Opens Door to Clinician-Style Tests for Medical AI

The agency is exploring a shift from attempting to test every possible AI interaction toward proving that a finished device can safely perform a defined clinical role—and continue doing so after changes.

By 3 min read
FDA Opens Door to Clinician-Style Tests for Medical AI

Listen to this story

The audio brief

About 1:34
0:001:34
Read transcript
The FDA is considering a new way to test medical AI: prove that a finished product can competently perform a defined clinical job, rather than trying to predict every prompt and interaction it might encounter. The idea appears in an exploratory discussion paper released August 18. It is not draft or final guidance, so it creates no new obligations yet. The proposed assessment would focus on the device clinicians and patients actually use—including its prompts, interface, safeguards, and workflow—not just the underlying foundation model. Evidence could come in three layers. Nonclinical benchmarks might test medical knowledge, reasoning, safety behavior, communication, and performance across patient groups and operating conditions. Clinical confirmation would examine how the device performs with real patients, clinicians, and workflows. The amount of evidence would depend on the intended use and the potential harm of a wrong answer; a prospective trial would not automatically be required. After deployment, manufacturers might face periodic retesting, clinician review of outputs, and monitoring for performance degradation. Changes to the software or a third-party foundation model could trigger another assessment. The proposal covers conversational tools and agentic systems that plan and execute multistep tasks. The FDA is collecting feedback through October 19, 2026. The key unresolved question is whether clinician-style competency testing can become a workable substitute for exhaustive testing of generative AI’s open-ended behavior.

Story brief

3 key points

The FDA is testing a regulatory model that could let generative-AI medical-device makers prove competence on a defined clinical job instead of exhaustively testing every possible prompt. Its August 18 discussion paper proposes evidence spanning nonclinical benchmarks, clinical confirmation, and post-market surveillance, with reassessment potentially triggered by software or foundation-model changes. The proposal...

  1. 01

    The agency would assess the finished device—including prompts, interface, safeguards, and workflow—not the underlying foundation model alone.

  2. 02

    Clinical evidence would depend on intended use and potential harm; a prospective trial may not be required for every device.

  3. 03

    Possible post-deployment controls include periodic retesting, clinician review of outputs, and monitoring for performance degradation.

Generative AI medical devices may eventually be judged less by whether regulators can anticipate every possible prompt and more by whether a finished product can reliably carry out its intended clinical task. The FDA has opened a public discussion on that possibility, including clinician-style competency testing, clinical confirmation, and monitoring after deployment.

The agency released Considerations for the Regulation of Generative AI-Enabled Medical Devices on August 18. It is seeking feedback through October 19, 2026, but the paper is neither draft nor final guidance. That distinction matters: the document starts a regulatory-design process rather than imposing new requirements on developers or healthcare providers.

The problem is that generative systems do not behave like conventional medical software with a defined function and a predictable range of inputs and outputs. They can take open-ended instructions, return different answers to similar prompts, and change when their underlying model, safeguards, or data sources change. Testing every possible interaction may therefore be impractical.

The FDA is instead considering an approach in which a manufacturer demonstrates that the completed device competently performs its intended clinical task. The assessment would focus on the product patients and clinicians actually use, rather than evaluating a foundation model in isolation. One general-purpose model can support products with different prompts, interfaces, safeguards, and clinical uses, making that distinction consequential.

A possible evidence path has three parts

  • Nonclinical benchmarking could examine clinical knowledge, analytical ability, safety behavior, communication, and performance across patient groups and operating conditions.
  • Clinical confirmation would test whether the device works with patients, clinicians, and real workflows. The needed evidence would vary with intended use and the potential harm of an incorrect output; a prospective clinical trial would not necessarily be required in every case.
  • Post-deployment oversight could include periodic retesting, clinician review of real-world outputs, and monitoring for performance deterioration. Relevant software or third-party foundation-model changes could trigger another assessment.

The paper also reaches beyond systems that generate text or other outputs in a single exchange. It covers foundation models and agentic systems that can plan and execute multistep tasks, alongside premarket testing and risk-proportionate postmarket monitoring. The effort is led by the FDA’s Digital Health Center of Excellence within the Center for Devices and Radiological Health.

The clinician analogy gives the proposal a practical frame: regulators can test underlying knowledge, observe performance in clinical settings, and maintain oversight without predicting every situation a physician will face. The FDA is examining whether a version of that model can work for AI while accounting for the technical and legal differences of regulated devices.

The unresolved work is substantial. The paper identifies possible approaches and asks targeted questions intended to shape a future framework; it does not settle what evidence will be sufficient for a given device, when reassessment will be required, or which proposals will become policy. Stakeholders now have until October 19 to contest, refine, or support the model before the FDA decides its next step.

Sources

  1. pymnts.comFDA Weighs Clinician-Style Tests for Generative AI Medical Devices | PYMNTS.com
  2. ascopost.comFDA Seeks Public Input Relating to Regulatory Considerations for Generative AI–Enabled Medical Devices