25 August 2026 · 3 min read
GenAI in Regulated Healthcare: The FDA and EMA Draw the Rules
Regulators on both sides of the Atlantic are beginning to define how generative AI earns its place in regulated healthcare — and I think the performance-based approach both the FDA and EMA are circling is exactly the right frame.
Author
TL;DR
The FDA and EMA are each moving to set rules for generative AI in regulated healthcare, from medical device software to medicines manufacturing. EMA's draft Annex 22 covers traditional, static ML models and currently keeps dynamic and generative models out of critical use — though EMA is already working on how to bring them in. The FDA's discussion paper proposes a competency-based approach: benchmark what a system knows and can do, then confirm it clinically, including for agentic systems. What links both efforts is a shared instinct to judge these systems by the performance they can demonstrate for a defined use — an approach I think rewards rigorous validation. The critical next step is defining what good evidence looks like for nondeterministic outputs: which competencies matter, how to test them, and where human oversight must remain.
The next phase of generative AI regulation is taking shape on both sides of the Atlantic.
On 18 August the FDA opened a discussion on how it might evaluate generative AI-enabled medical devices. In parallel, Europe is building draft Annex 22, which keeps generative models out of another regulated area, critical manufacturing, for now.
What EMA's Draft Annex 22 Covers — and Leaves Out
EMA's draft Annex 22 would set the rules for AI in medicines manufacturing. It covers the "traditional" machine learning pharma has used for years: a model trained on data, then held static in use, giving the same answer to the same input. Predictable, and checkable.
Dynamic, probabilistic and generative models, i.e pretty much everything we consider GenAI/agentic today, sit outside the scope for critical use. EMA is already considering how to bring them in. It held a two-day workshop this summer on the safeguards that might make it possible.
The FDA's Competency-Based Alternative
The FDA's device center took a different route. Its discussion paper considers a competency-based approach: evaluate the system the way you would credential a doctor. Benchmark what it knows and can do, then confirm it clinically. The paper even takes on agentic systems directly, asking how to handle the reduced room for human review.
The Idea Linking Both Approaches
These are different domains. Annex 22 covers the factory floor, while the FDA paper covers medical device software used in patient care. What links them is one idea both are circling: judge these systems by the performance they can demonstrate for a defined use. I really like that approach. It rewards the teams doing the hard validation work, and it gives newer models a real path into regulated settings.
Regulators have a great opportunity here. The next step is to define what good evidence looks like when outputs can vary in nondeterministic systems like GenAI: which competencies matter, how they should be tested, and where human oversight remains essential. Get that right, and newer systems have a credible path into regulated use without lowering the bar on safety.
Key takeaways
- Both the FDA and EMA are actively shaping how generative AI will be evaluated in regulated settings — this is no longer a future question.
- EMA's draft Annex 22 currently limits regulated use to traditional, static ML models; dynamic and generative models are explicitly out of scope for critical use for now.
- EMA held a two-day workshop this summer on the safeguards that could eventually bring GenAI and agentic models into scope.
- The FDA's competency-based approach treats a GenAI system the way you would credential a doctor — benchmark it, then confirm it clinically.
- The FDA paper directly addresses agentic systems and the reduced room for human review they create.
- I believe a performance-based standard rewards teams doing hard validation work and gives newer models a credible path into regulated settings.
- The defining challenge ahead is specifying what good evidence looks like for nondeterministic outputs — which competencies matter, how to test them, and where human oversight stays essential.
Related insights
9 Oct 2024
A great talk from Gilead's Jeremy Zhang on scaling LLMs like OpenAI and Claude to the clinical po...
14 Jul 2026
Two hits, two partials, one miss
5 Mar 2026
FDA GenAI clinical care regulation: why Breakthrough Device Designation matters
16 Sept 2025
Healthcare AI agents: why evaluation must move from answers to actions