
EU GMP Annex 22.
Annex 22 is the planned AI annex to EudraLex Volume 4, the EU GMP guideline. It is the first standalone set of rules for AI models inside the computerized systems used in pharmaceutical manufacturing. The computerized system is still validated under Annex 11, and the embedded AI model faces requirements that go further. The text is still a draft, yet it can already shape how inspectors proceed.
Ten sections, one narrow boundary for critical applications.
The draft applies to computerized systems in the GMP environment in which AI models operate in critical applications. Critical here means a direct impact on patient safety, product quality or data integrity. Such models typically predict values or classify data. We read the scope broadly, from production to quality processes such as deviation management. The draft covers machine-learning models that acquired their function through training.
In our reading of the ongoing discussion, "critical" is not a matter of direct impact alone. Errors that are hard to detect and a high degree of automation also point to a higher classification. The text deliberately draws a tight boundary and, for critical applications, provides only for static models with deterministic output. The ten sections show how this fits into the wider picture of AI in GxP.
Scope & principles · Sections 1-2
Section 1 limits the scope to critical applications and to static models with deterministic output. Section 2 requires the parties to work closely together (process SMEs, QA, data science, IT and consultants), each with defined responsibilities. The regulated user reviews the related documentation, including for models supplied by a vendor. How far the measures go depends on the risk.
Intended Use · Section 3
The draft expects a detailed description of the model's task, with input data (input sample space), limitations and possible bias. The process SME is responsible for that description, which is to be approved before testing. The draft applies "Intended Use" to the model. In our view the purpose comes from the process, and the model performs a task within it.
Acceptance criteria · Section 4
Test metrics and acceptance criteria are set before testing. The acceptance criteria should be at least as high as the performance of the process the model replaces.
Test data · Sections 5-6
Test data should be representative and stratified, and it is to be kept strictly separate from model development, technically and in terms of personnel, secured by access control and an audit trail. Where separation of personnel is not possible, the draft calls for the four-eyes principle.
Testing & explainability · Sections 7-9
Testing runs under an approved test plan and checks whether the model generalizes. During testing in critical applications, the system records which features of the test data contributed to the outcome, and approval of the test results includes a review of whether the model relies on relevant features. Techniques such as feature attribution appear in the draft only as examples ("where applicable"). The model can flag uncertain results as "undecided". How they are handled is, in our view, for the process to decide.
Operation · Section 10
Before go-live, the draft puts the model, the system and the process under change control, with the tested model also under configuration control. Model performance and input data need monitoring for drift, and where the testing effort for the model has been reduced, records have to be kept. Depending on criticality, this may mean that a human reviews or tests every output according to a defined procedure.
The thread running through these sections is a shift in evidence. Statistical evidence takes the place of the single deterministic test. Test data comes with its own requirements, and metrics show how well the model performs. For GMP this way of thinking is not new. Since 2015, Annex 15 has required ongoing process verification across the lifecycle. Where appropriate, statistical tools support the conclusions on process variability and capability (5.29 to 5.32). Read that way, Annex 22 carries that evidence over from processes to models.
Rule-based automation without machine learning is purely an Annex 11 case. Static machine-learning models are the subject of Annex 22. For critical applications, by contrast, the draft does not provide for dynamic, probabilistic or generative models. Outside critical applications they can still be used, with the computerized system they run in subject to the risk assessment under the current Annex 11. In our view, they should always run under human oversight. The draft makes such oversight an explicit condition for generative AI and LLMs.
Where the text stands today.
1,359 comments from 79 organizations in 39 countries, submitted by industry, regulators and academia, reflect the interest in the draft. The road to the final text runs through five stages.
- 1
Draft and consultation
In July 2025, the European Commission put the draft of the first standalone AI annex to the EU GMP guideline out for consultation. It was prepared by the Annex 22 drafting group of the GMDP Inspectors Working Group. The comment period ran until October 2025. The annex belongs to the EudraLex Volume 4 revision package, together with Chapter 4 (Documentation) and the Annex 11 revision.
- 2
Evaluation of comments
The evaluation presented by the chair of the GMDP Inspectors Working Group in April 2026 shows clear themes, above all the call for a risk-based approach without categorical technology exclusions.
- 3
Expert workshop
Day one, 30 June 2026, was an open session in which experts nominated by the associations presented their views and evidence. Day two, 1 July, was a closed session without broadcast. The six topic areas ranged from regulatory pathways for adaptive models through human oversight to cybersecurity.
- 4
Report and revised draft
A report with the expert contributions has been announced as the next step. After that, a version that incorporates the outcomes of the consultation and the workshop is expected. Our overview Three Signals, One Line shows how closely the signals from Barcelona (rapporteur of the Annex 22 drafting group), Boston (FDA expert) and the EMA workshop align. There is a transatlantic signal as well. On 14 January 2026, EMA and FDA jointly published Guiding Principles of Good AI Practice. Principle 2 ties validation, risk mitigation and oversight to the context of use and the model risk that has been determined. Principle 9 calls for scheduled monitoring and periodic re-evaluation, for example to catch data drift.
- 5
Finalization
Publication and transition periods are still pending. Until then Annex 11 applies unchanged.
Do generative AI and LLMs stay out?
Yes. The draft does not apply to generative AI and Large Language Models, and in critical GMP applications they should not be used. Outside critical applications they can still be used, always with human review of the outputs (Human-in-the-Loop, HITL). Responsibility for their use lies with the regulated company. The EMA summarized the key themes from the consultation as follows:
What will decide the final text of Annex 22?
The consultation and the expert workshop have narrowed the debate to three questions, and their answers will determine how far the final text departs from the consultation draft. The first question is whether the boundary stays tied to the model class or moves to the risk of the application. The second is about oversight: a fixed human-in-the-loop mechanism or human oversight with several forms. The third targets the evidence the EMA will require for guardrails before dynamic models are admitted to GMP applications. So far, only the direction is public. According to its workshop announcement, the EMA is examining guardrails and other control and mitigation measures.
Does the boundary stay tied to the model class?
For critical applications, the draft excludes dynamic, probabilistic and generative models. The logic of quality risk management points the other way. Higher risk means stronger controls, not a blanket ban. The EMA named the first workshop topic area "Regulatory Pathways for Adaptive/probabilistic AI models in GMP". Its guiding question, in the EMA's own words, was "How could we accommodate adaptive and probabilistic models in Annex 22?" The debate is therefore about a regulatory pathway for precisely these models. In its announcement of the workshop, the EMA counts guardrails and other control measures among the elements of a proposed risk-based approach.
How much oversight will the final text require?
The draft relies on Human-in-the-Loop in three places: in the scope as the condition for generative models outside critical applications, and in sections 3.3 and 10.5 as compensation for reduced testing effort. The consultation mainly asked for clarity here. At the expert workshop, EFPIA took the position that oversight is a concept with several forms, such as HITL, monitoring in operation or governance controls. On that reading, the appropriate form follows from the risk assessment. The shift from mechanism to concept would be one of the clearest changes of direction in the final rules. Our contribution to the GAMP D-A-CH Forum argues why that shift begins with responsibility: Who Pushes Back When the System Speaks? For ISPE iSpeak we made the case that a human in the loop does not by itself establish control: Human-in-the-Loop as an Illusion of Control?
What evidence for guardrails?
The workshop explicitly asked for evidence: how effective are the control measures and guardrails meant to support a risk-based approach? On the first workshop day, the PDA presented five GMP use cases, from adaptive cleanroom climate control and AI agents in deviation management to model-predictive control of a bioreactor. A three-layer guardrail architecture was presented alongside. Outlier detection sits at the input, confidence methods inside the model, and hard rule limits with statistical drift detection using CUSUM and EWMA at the output. The ISPE added that guardrails are themselves classifying systems with their own error rates. They therefore need the same evidence as the model. We expect any control that is supposed to lower a model's risk to carry its own evidence of effectiveness, with metrics, a testing approach and monitoring. Without it, the measure merely shifts the risk from the model into an untested control.
How do Chapter 4, Annex 11 and Annex 22 interlock?
The three drafts of the EudraLex Volume 4 revision are interrelated. Chapter 4 governs documentation. It leaves responsibility for the integrity of documents, records and data with the regulated user, even when AI or other automated means produce or process that content. Where AI supports a decision in manufacturing, Annex 11 applies. In critical applications, Annex 22 applies on top of it. The Annex 11 draft forms the framework for every computerized system, with risk management, security, access management, audit trails and periodic review. That includes the principle that a new system should not increase the overall risk of the process. Annex 22 builds on that and, as supplementary guidance, describes what Annex 11 does not cover for a trained model: intended use of the model as the draft understands it, test data kept separate from the training and validation data used in model development, acceptance criteria at least at the level of the process it replaces, explainability, confidence thresholds and monitoring in operation.
For the principle of "no decrease in performance", Annex 22 refers to Annex 11 clause 2.7. In the Annex 11 draft that principle sits under 2.8, while clause 2.7 there covers security.
The draft draws the line clearly.
"Following the above, the document does not apply to Generative AI and Large Language Models (LLM), and such models should not be used in critical GMP applications."
EU GMP Annex 22 (Artificial Intelligence), draft for consultation 2025, Section 1 "Scope".
Whether this exclusion holds in the final text remains an open question. The logic of evidence behind it applies either way: intended use, independent test data, monitoring and a criticality assessment that looks beyond direct impact. In our view the draft demands more than necessary in section 6.5, staff independency, for example. The intent behind it is clear. Test data must not quietly flow back into training just so that the model passes on paper. The draft's own answer, separating staff or, failing that, applying the four-eyes principle, is one way of doing it. Depending on an organization's digital maturity, more effective solutions exist: access control and an integrated MLOps infrastructure with data lineage that makes it traceable which data were used where and for which model variant. The review then checks whether that infrastructure was used as required. That keeps the focus on the question that matters, whether the test data are suitable and meaningful enough to allow a realistic performance statement for operation. In a similar way, 4.3 and 8.2 mix the what with the how. In 4.3, the principle of "no decrease in performance" presupposes a known predecessor process and a one-to-one replacement. In 8.2, the draft has image-processing models in mind, whose features can be reviewed. For other model classes such a review is hardly possible in that sense any more, and interpretability appropriate to the context of use would be the better goal.
What you can put in place before the final text.
Annex 11 provides the framework for the building blocks of the draft, and oversight belongs in AI governance with human oversight.
We worked through what this looks like in practice in June 2026, in a workshop at the ISPE AI in Life Sciences Summit in Boston. A RAG system built on a third-party language model classifies deviation reports in GMP operations as critical, major or minor. QA exercised human oversight and reviewed every classification before it took effect. Whether such an application counts as critical is decided by the operator's criticality assessment, independently of the downstream review. The supplier and the regulated company sorted ten gaps independently of each other into dealbreakers, fixable items and acceptable risks, then negotiated a joint list of mandatory evidence. The company could not delegate the intended use in its own process, the acceptance criteria or the residual risk. A model card, class-specific metrics and a test set kept separate from training were part of the evidence the company required from its supplier.
These are the building blocks we assess and put in place with you:
Frequently asked questions about Annex 22.
No date has been set. The consultation ended in October 2025, and the EMA expert workshop took place on 30 June and 1 July 2026. A workshop report has been announced. After that, a revised draft, finalization and transition periods are expected. Until then the computerized system is validated under Annex 11. Quality risk management under ICH Q9(R1) guides the extent of the evidence.
Formally, generative AI and LLMs fall outside the scope of Annex 22, yet the draft is clear about them. They should not be used in critical GMP applications, and the evidence sections apply only to static models. Outside critical applications they may be used, provided a human reviews the outputs. The regulated company is responsible for that use. This exclusion is the central point of criticism in the consultation, and the coming version may look different here.
No. The draft has no zones comparable to the cleanroom grades in Annex 1. It distinguishes along two axes. Applications count as critical or non-critical depending on their direct impact on patient safety, product quality or data integrity. Models are divided into static models with deterministic output and dynamic, probabilistic or generative models. Together, the two axes determine which sections apply and whether the draft provides for their use in critical applications.
No. Annex 22 is designed as supplementary guidance to Annex 11. The computerized system is still validated under Annex 11, and Annex 22 additionally governs the evidence for the embedded model, from intended use to operation.
Annex 22 concentrates on the model, that is, on training, testing and the evidence of its fitness for purpose. We take a wider view. An AI system or AI subsystem integrates one or more models and additional functions, such as controls ("guardrails"), into a system component. Performance can then be measured at several levels, so that the suitability of all subcomponents and their interplay can be checked.
Section 2.1 of the draft calls for close cooperation among everyone involved, from algorithm selection through training, validation (in the model-development sense) and testing to operation. It names process SMEs, QA, data science, IT and consultants, each with adequate qualifications, defined responsibilities and appropriate access rights. For generative models outside critical applications, qualified personnel carry the responsibility for ensuring the outputs are suitable for the intended use. The cooperation under section 2.1 is also Knowledge Management. ICH Q10 names it, alongside quality risk management, as one of the two enablers of the quality system. That keeps knowledge about models, data and their limits inside the organization.
Regulated companies can meet the core of the requirements before finalization. An AI inventory, intended use and criticality assessment per application, test metrics, independent test data, monitoring and human oversight can all be implemented under the current Annex 11. ICH Q9(R1) guides how far the evidence has to go. Since July 2025, the methodology for AI-enabled systems has had a home of its own, the ISPE GAMP Guide: Artificial Intelligence. The draft makes much of this explicit for AI models.
Last updated:
Annex 22 needs the full context.
Annex 22 readiness starts before the final text.
In an initial conversation, we explore what you plan to do with AI, where your applications stand today and what Annex 22 means for you. Free of charge, about 30 minutes.
Book an initial conversation


