
EU GMP Annex 22.
Annex 22 is the planned AI annex to EudraLex Volume 4, the EU GMP guideline. It is the first standalone set of rules for AI models inside the computerized systems used in pharmaceutical manufacturing. The computerized system remains validated under Annex 11; the AI model embedded in it faces additional requirements. The text is still a draft. Yet it can already shape how inspectors proceed.
Ten sections, one narrow boundary for critical applications.
The draft applies to computerized systems in the GMP environment wherever AI models operate in critical applications, that is, with a direct impact on patient safety, product quality or data integrity, for example by predicting or classifying data. That scope extends from production to quality processes such as deviation management. Criticality is not a matter of direct impact alone: it also rises where errors are hard to detect and the degree of automation is high. That is our reading of the ongoing discussion. The models covered are machine-learning models that acquire their function through training, and the draft draws a deliberately narrow boundary: in critical applications it provides only for static models with deterministic output. The ten sections show how all of this fits into the wider picture of AI in GxP.
Scope & principles · Sections 1-2
Section 1 sets the scope: critical applications, static models with deterministic output. Section 2 governs the division of labor: process SMEs, QA, data science and IT work closely together; documentation rests with the regulated user, including for supplier-provided models, and its extent follows the risk.
Intended Use · Section 3
The model's task is described in detail: input data (input sample space), limitations, possible bias. The process SME owns that description, and it is approved before testing. The draft transfers "Intended Use" to the model; in our view the purpose comes from the process, and the model performs a task within it.
Acceptance criteria · Section 4
Test metrics and acceptance criteria are fixed before testing. The model must perform at least as well as the process it replaces.
Test data · Sections 5-6
Representative, stratified test data, kept strictly separate from model development, both technically and organizationally: access control, audit trail, four-eyes principle.
Testing & explainability · Sections 7-9
An approved test plan demonstrating generalization, plus explainability with expert review; techniques such as feature attribution appear only as illustrations ("where applicable"). The model can flag uncertain results as "undecided"; the process defines how they are then handled.
Operation · Section 10
Change and configuration control before deployment; monitoring of model performance and input data for drift; human review under a defined procedure wherever it compensates for reduced testing effort.
One thread runs through all of these sections: what counts as evidence shifts from deterministic individual tests to statistical evidence. Test data must meet requirements of its own, and metrics have to hold up as the basis for control and effectiveness. That is how data scientists think, and it is also the core of quality risk management under ICH Q9(R1): deciding under uncertainty. Not a break with GxP logic, but its consistent application.
For perspective: rule-based automation without machine learning remains purely an Annex 11 case, while static ML models with deterministic output are Annex 22 territory. In critical applications the draft does not provide for dynamic, probabilistic or generative models. Everywhere else they remain usable, subject to the risk assessment the current Annex 11 requires of any computerized system, and in our view always under human oversight. The draft makes such oversight an explicit condition for generative AI and LLMs.
Where the text stands today.
The public consultation put a number on the industry's interest: it drew 1,359 comments from 79 organizations in 39 countries. The road to the final text in five stages:
- 1
Draft and consultation
July to October 2025: the EMA's GMDP Inspectors Working Group opens the first standalone AI annex to the EU GMP guideline for comment, as part of a EudraLex Volume 4 package that also revises Chapter 4 (Documentation) and Annex 11.
- 2
Evaluation of comments
The GMDP IWG evaluates the submissions. The EMA's summary shows clear priorities, above all the call for a risk-based approach instead of categorical technology exclusions.
- 3
Expert workshop
30 June and 1 July 2026: day one is an open session in which association-nominated experts present their views and evidence; day two, a closed session in which the EMA's Annex 22 drafting group works through the contributions. The six topic areas range from regulatory pathways for adaptive models through human oversight to cybersecurity, with one goal: evidence on the control measures and guardrails a risk-based approach would need.
- 4
Report and revised draft
A report capturing the expert input has been announced; a version that incorporates the outcomes of the consultation and the workshop is expected to follow. How closely the signals on the Annex 22 draft from Barcelona, Boston and the EMA workshop already align is what our overview Three Signals, One Line shows.
- 5
Finalization
Publication and transition periods are still to come. Until then, Annex 11 applies; ICH Q9(R1) provides guidance for the risk-based approach.
Do generative AI and LLMs stay out?
The draft says yes. It does not apply to generative AI and Large Language Models, and they should not be used in critical GMP applications; outside critical applications, their use remains possible, always with human review of the outputs (Human-in-the-Loop, HITL). Responsibility for their use remains with the regulated company. The consultation's main comments, as summarized by the EMA, start exactly here:
What will decide the final text of Annex 22?
The consultation and the expert workshop have condensed the debate into three questions. They will decide how far the final text departs from the draft:
Risk principle or model-class boundary?
For critical applications, the draft excludes dynamic, probabilistic and generative models. The logic of quality risk management argues for the opposite approach: higher risk means stronger controls, not a blanket ban. The EMA itself named the first topic area of its expert workshop "Regulatory Pathways for Adaptive/probabilistic AI models in GMP": what is on the table is a regulatory pathway for precisely these models, no longer just the line that keeps them out.
HITL mechanism or human oversight?
The draft anchors Human-in-the-Loop in three places: in the scope as the condition for generative models outside critical applications, and in sections 3.3 and 10.5 as compensation for reduced testing effort. The consultation pushes back: oversight is a concept with several forms, and the appropriate form follows from the risk assessment. The shift from mechanism to concept would be one of the clearest changes of direction in the final text. Why this change of direction starts with responsibility is the case our position from the GAMP D-A-CH Forum makes: Who Pushes Back When the System Speaks? Why a human in the loop does not by itself establish control is the argument we make for ISPE iSpeak: Human-in-the-Loop as an Illusion of Control?
What evidence for guardrails?
The workshop asked explicitly for evidence: how effective are the control measures and guardrails meant to support a risk-based approach? Our position: a control intended to reduce a model's risk needs its own evidence of effectiveness, with metrics, a testing approach and monitoring. Without it, the risk simply moves from the model into an untested control.
The draft draws the line clearly.
"Following the above, the document does not apply to Generative AI and Large Language Models (LLM), and such models should not be used in critical GMP applications."
EU GMP Annex 22 (Artificial Intelligence), draft for consultation 2025, Section 1 "Scope".
Whether this line holds in the final text remains an open question. The logic of evidence behind it stands either way: intended use, independent test data, monitoring and a criticality assessment that weighs not only direct impact but also the detectability of errors and the degree of automation. Put that in place today and you are prepared for either outcome.
Preparation does not wait for a final text.
The draft is a checklist in all but name: the core of its requirements can be established today with Annex 11 and ICH Q9(R1), embedded in AI governance with human oversight. These are the building blocks we assess and put in place:
Frequently asked questions about Annex 22.
No date is set. The consultation closed in October 2025 and the EMA's two-day expert workshop took place on 30 June and 1 July 2026; a workshop report has been announced, with a revised draft expected next, followed by finalization and transition periods. Until then, the computerized system is validated under Annex 11; quality risk management under ICH Q9(R1) guides the depth of evidence.
Formally, generative AI and LLMs fall outside its scope. The draft nevertheless makes a clear statement about them: they should not be used in critical GMP applications, and the evidence sections apply only to static models with deterministic output. Outside critical applications their use remains possible, always with human review of the outputs, and responsibility stays with the regulated company. This very boundary is the central point of contention in the consultation, and the final text may read differently here.
No. Annex 22 is designed as supplementary guidance to Annex 11: the computerized system remains validated under Annex 11; Annex 22 additionally governs the evidence for the embedded model, from intended use to operation.
Annex 22 focuses on the model: training, testing and the evidence of its fitness for purpose. We take a wider view: an AI system or AI subsystem brings together one or more models and additional functions, such as controls (often called "guardrails"), into a single system component. Performance can then be measured at several levels, so the fitness for purpose of each subcomponent, and of their interplay, can be assessed.
Section 2.1 of the draft calls for close cooperation between everyone involved, from algorithm selection through training, validation and testing to operation. It names process SMEs, QA, data scientists, IT and external consultants, each with adequate qualifications, defined responsibilities and an appropriate level of access. Where generative models run outside critical applications, qualified personnel remain responsible for ensuring the outputs are suitable for the intended use. This cooperation is also Knowledge Management in the sense of ICH Q10, where it stands alongside quality risk management as one of the two enablers of the quality system: knowledge about models, data and their limits stays with the organization, not with individuals.
Everything that matters: inventory the AI applications, document intended use and criticality, define test metrics, secure independent test data, establish monitoring and human oversight. Every one of these building blocks can be implemented today under Annex 11, with ICH Q9(R1) informing how deep the evidence needs to go. The draft makes much of this explicit for AI models and in places demands more.
Last updated:
Annex 22 needs the full context.
Annex 22 readiness starts before the final text.
In an initial conversation, we explore what you plan to do with AI, where your applications stand today and what Annex 22 means for you. Free of charge, about 30 minutes.
Book an initial conversation


