
EU GMP Annex 22.
Annex 22 is the planned AI annex to EudraLex Volume 4, the EU GMP guideline - the first standalone set of rules for AI models inside the computerized systems used in pharmaceutical manufacturing. The computerized system remains validated under Annex 11; the AI model embedded in it faces additional requirements. The text is still a draft - yet it can already shape how inspectors approach AI.
Ten sections - one narrow boundary for critical applications.
The draft applies to computerized systems across the GMP environment - from production to quality processes such as deviation management - wherever AI models operate in critical applications, meaning they bear directly on patient safety, product quality or data integrity, for example by predicting or classifying data. Criticality is not a matter of direct impact alone: it also rises where errors are hard to detect and the degree of automation is high - our reading of the ongoing discussion. The models covered are machine-learning models that acquire their function through training, and the draft draws a deliberately narrow boundary: in critical applications, only static models with deterministic output are in scope. The ten sections show how all of this fits into the wider picture of AI in GxP.
Scope & principles · Sections 1-2
Section 1 sets the scope: critical applications, static models with deterministic output. Section 2 governs the division of labor: process SMEs, QA, data science and IT work closely together; documentation rests with the regulated user - including for supplier-provided models, and how far it has to go follows the risk.
Intended use · Section 3
The model's task is described in detail - input data (input sample space), limitations, possible bias - owned by the process SME and approved before testing. The draft transfers "Intended Use" to the model; in our view the purpose comes from the process, and the model performs a task within it.
Acceptance criteria · Section 4
Test metrics and acceptance criteria are fixed before testing. The model must perform at least as well as the process it replaces.
Test data · Sections 5-6
Representative, stratified test data - kept strictly separate from model development, both technically and organizationally: access control, audit trail, four-eyes principle.
Testing & explainability · Sections 7-9
An approved test plan demonstrating generalization; explainability with expert review - the draft names techniques such as feature attribution only as illustrations ("where applicable"); and the model can flag uncertain results as "undecided" instead of forcing an unreliable classification - how they are then handled is defined in the process.
Operation · Section 10
Change and configuration control before deployment; monitoring of model performance and input data for drift; human review under a defined procedure wherever it compensates for reduced testing effort.
What runs through all of these sections is a change in the form of the evidence: from deterministic single tests to statistical evidence. Test data that must itself meet requirements, metrics robust enough to carry claims about control and effectiveness - this is how data scientists think, and it is also the core of quality risk management under ICH Q9(R1): deciding under uncertainty. Not a break with GxP logic, but its consistent application.
For perspective: rule-based automation without machine learning remains a straightforward Annex 11 case, while static ML models with deterministic output are Annex 22 territory. In critical applications the draft leaves no room for dynamic, probabilistic or generative models. Everywhere else they remain usable - with the risk assessment the current Annex 11 requires of any computerized system, and in our view always under human oversight, which the draft makes an explicit condition for generative AI and LLMs.
Where the text stands today.
The public consultation put a number on the industry's interest: it drew 1,359 comments from 79 organizations in 39 countries. The road to the final text in five stations:
- 1
Draft and consultation
July to October 2025: the EMA's GMDP Inspectors Working Group opens the first standalone AI annex to the EU GMP guideline for comment - as part of a EudraLex Volume 4 package that also revises Chapter 4 (Documentation) and Annex 11.
- 2
Evaluation of comments
The GMDP IWG works through the submissions - with clear priorities per the EMA's summary, above all the call for a risk-based approach instead of categorical technology exclusions.
- 3
Expert workshop
30 June and 1 July 2026: day one an open session in which association-nominated experts present their views and evidence; day two a closed session in which the EMA's Annex 22 drafting group works through the contributions. Six topic areas span regulatory pathways for adaptive models, human oversight and cybersecurity, with one goal: evidence on the control measures and guardrails a risk-based approach would need.
- 4
Report and revised draft
A report capturing the expert input has been announced; expected after that is a version that works in the consultation and workshop outcomes.
- 5
Finalization
Publication and transition periods are still to come. Until then, Annex 11 applies - with ICH Q9(R1) providing guidance for the risk-based approach.
Do generative AI and LLMs stay out?
The draft says yes. It does not apply to generative AI and Large Language Models, and they should not be used in critical GMP applications; outside critical applications, their use remains possible - always with human review of the outputs (Human-in-the-Loop, HITL). Responsibility for their use remains with the regulated company. This is precisely where the consultation's comment priorities, as summarized by the EMA, take issue:
Our reading: the logic behind this critique is that of quality risk management - higher risk means stronger controls, not a blanket ban. The EMA itself has put this direction on the agenda of its expert workshop: the first of six topic areas ("Regulatory pathways for adaptive / probabilistic AI models in GMP") asks outright, "How could we accommodate adaptive and probabilistic models in Annex 22?" What is being negotiated is a pathway for these models, no longer just their exclusion. Whether the final text follows that line is for the ongoing revision to decide.
The draft draws the line clearly.
"Following the above, the document does not apply to Generative AI and Large Language Models (LLM), and such models should not be used in critical GMP applications."
EU GMP Annex 22 (Artificial Intelligence), draft for consultation 2025, Section 1 "Scope".
Whether this line holds in the final text remains an open question. The logic of evidence behind it holds either way: intended use, independent test data, monitoring - and a criticality assessment that weighs not only direct impact but also how detectable errors are and how far the process is automated. Put that in place today and you are prepared for either outcome.
Preparation does not wait for a final text.
The draft is a checklist in raw form: every one of its requirements can be put in place today with Annex 11 and ICH Q9(R1), embedded in AI governance with credible human oversight. These are the building blocks we examine and establish:
Frequently asked questions about Annex 22.
No date is set. The consultation closed in October 2025 and the EMA's two-day expert workshop took place on 30 June and 1 July 2026; a workshop report has been announced, with a revised draft expected next, followed by finalization and transition periods. Until then, the computerized system is validated under Annex 11 - with quality risk management per ICH Q9(R1) guiding the depth of evidence.
Formally they fall outside its scope - yet it is precisely about them that the draft makes its clearest statement: they should not be used in critical GMP applications, and the evidence sections apply only to static models with deterministic output. Outside critical applications their use remains possible - always with human review of the outputs, and responsibility stays with the regulated company. This very boundary is the central point of contention in the consultation, and the final text may well read differently here.
No. Annex 22 is designed as supplementary guidance to Annex 11: the computerized system remains validated under Annex 11; Annex 22 additionally governs the evidence for the embedded model - from intended use to operation.
Annex 22 focuses on the model - training, testing and the evidence of its fitness for purpose. We take a wider view: an AI system or AI subsystem brings together one or more models and additional functions, such as controls (often called "guardrails"), into a single system component. Performance can then be measured at several levels, which means verifying the fitness for purpose of every subcomponent and of the way they work together.
Everything that matters: inventory the AI applications, document intended use and criticality, define test metrics, secure independent test data, establish monitoring and human oversight. Every one of these building blocks can be put in place today under the current Annex 11, with ICH Q9(R1) guiding the depth of evidence - the draft merely makes them explicit for AI models.
Annex 22 needs the full context.
Annex 22 readiness starts before the final text.
In an initial consultation, we explore what you are planning with AI, where your applications stand today and what Annex 22 means for you. Free of charge, about 30 minutes.
Book an initial consultation


