EU GMP Annex 22 - the new AI annex to the EU GMP guideline for AI models in pharmaceutical manufacturing
QFINITY · EU GMP Annex 22

EU GMP Annex 22.

Annex 22 is the planned AI annex to EudraLex Volume 4, the EU GMP guideline - the first standalone set of rules for AI models in computerized systems in pharmaceutical manufacturing. The computerized system continues to be validated under Annex 11; the embedded AI model is subject to further-reaching requirements. The text is still a draft. Yet it can already influence how inspectors approach AI.

The draft

Ten sections - one narrow boundary for critical applications.

The draft applies to computerized systems across the GMP environment - from production to quality processes such as deviation management - where AI models operate in critical applications: with direct impact on patient safety, product quality or data integrity, for example to classify data. Critical means more than direct impact alone: where errors are hard to detect and the degree of automation is high, criticality rises - our reading of the ongoing discussion. The draft covers machine-learning models that have obtained their functionality through training; and it draws a deliberately narrow boundary: in critical applications, only static models with deterministic output are within scope. How this fits into the bigger picture of AI in GxP becomes clear from the ten sections.

Scope & principles · Sections 1-2

Section 1 sets the scope: critical applications, static models with deterministic output. Section 2 governs the division of labor: process SMEs, QA, data science and IT work closely together; documentation rests with the regulated user - including for supplier-provided models, with the extent following the risk.

Intended use · Section 3

The model's task is described in detail - input data (input sample space), limitations, possible bias; owned by the process SME, approved before testing. The draft transfers "Intended Use" to the model - in our view, the purpose comes from the process; the model fulfills a task within it.

Acceptance criteria · Section 4

Test metrics and acceptance criteria are fixed before testing. Model performance must not fall below the process the model replaces.

Test data · Sections 5-6

Representative, stratified test data - strictly separated from model development, technically and organizationally: access control, audit trail, four-eyes principle.

Testing & explainability · Sections 7-9

An approved test plan demonstrating generalization; explainability with expert review - the draft names techniques such as feature attribution only as illustrations ("where applicable"); uncertain results count as "undecided" - and then require defined handling in the process.

Operation · Section 10

Change and configuration control before deployment; monitoring of model performance and input data for drift; human review under a defined procedure where it compensates for reduced testing effort.

For perspective: rule-based automation without machine learning remains a straightforward Annex 11 case. Static ML models with deterministic output are Annex 22 territory. For critical applications, the draft does not admit dynamic, probabilistic or generative models - outside of those, they remain usable, with risk assessment and human oversight.

Static models Deterministic output Critical GMP applications
Roadmap

Where the text stands today.

The public consultation made the industry's interest measurable: 1,359 comments from 79 organizations in 39 countries were submitted. The road to the final text in five stations:

  1. 1
    Draft and consultation

    July to October 2025: the EMA's GMDP Inspectors Working Group puts the first standalone AI annex to the EU GMP guideline out for comment - packaged with the EudraLex Volume 4 revision of Chapter 4 (Documentation) and Annex 11.

  2. 2
    Evaluation of comments

    The GMDP IWG works through the submissions - with clear priorities per the EMA's summary, above all the call for a risk-based approach instead of categorical technology exclusions.

  3. 3
    Expert workshop

    Late June 2026: the EMA explores with experts nominated by the industry associations whether risk-based controls can keep bias and hallucination under control in critical applications.

  4. 4
    Revised draft

    Expected as the next step: a version that works in the consultation and workshop outcomes.

  5. 5
    Finalization

    Publication and transition periods are still to come. Until then, Annex 11 applies - with ICH Q9(R1) providing guidance for the risk-based approach.

The contested question

Do generative AI and LLMs stay out?

The draft says yes. It does not apply to generative AI and Large Language Models, and they should not be used in critical GMP applications; outside critical applications, their use remains possible - always with human review of the outputs (Human-in-the-Loop, HITL). Responsibility for their use remains with the regulated company. This is precisely where the consultation's comment priorities, as summarized by the EMA, take aim:

  • Risk instead of technology bans

    Assessment per use case instead of categorical exclusion of entire model classes.

  • Define "critical application"

    A robust criticality methodology instead of ad-hoc classification.

  • Harmonization

    Coherence with ICH Q9, Annex 11 and the EU AI Act instead of parallel control worlds.

  • Training data and bias

    Dedicated requirements for provenance, representativeness and bias assessment.

  • Clarify human oversight

    Design oversight based on risk - HITL is one form of it, not the only one.

  • Model lifecycle

    Rules for versioning, retraining and AI functions delivered from the cloud.

Our reading: the logic behind this critique is that of quality risk management - higher risk means stronger controls, not a blanket ban. Whether the final text follows this line is up to the ongoing revision.

Annex 22 makes explicit much of what should already be good practice today - though in places it demands more than necessary.
The wording

The draft draws the line clearly.

"Following the above, the document does not apply to Generative AI and Large Language Models (LLM), and such models should not be used in critical GMP applications."

EU GMP Annex 22 (Artificial Intelligence), draft for consultation 2025, Section 1 "Scope".

Our position

Whether this line holds in the final text is open. The logic of evidence behind it applies either way: intended use, independent test data, monitoring - and a criticality assessment that considers not only direct impact but also the detectability of errors and the degree of automation. Whoever establishes it today is prepared for both outcomes.

Readiness

Preparation does not wait for a final text.

The draft is a checklist in raw form: every one of its requirements can be established today with Annex 11 and ICH Q9(R1), embedded in AI governance with credible human oversight. The building blocks we examine and put in place:

AI inventory: where do models affect patient safety, product quality or data integrity?
Purpose documented per model and verified against the specification - input data, limitations
Criticality assessed: decision consequence and the model's influence on the process
Test metrics and acceptance criteria defined and approved before testing
Test data representative, stratified and demonstrably independent of development
Explainability evidence for critical applications - method suited to the model
Confidence thresholds with a defined way of handling uncertain results
Operation under control: change control, drift monitoring, human oversight
FAQ

Frequently asked questions about Annex 22.

No date is set. The consultation closed in October 2025 and the EMA's expert workshop took place in late June 2026; a revised draft is expected next, followed by finalization and transition periods. Until then, the computerized system is validated under Annex 11 - with quality risk management per ICH Q9(R1) guiding the depth of evidence.

Formally they are outside its scope - yet the draft makes a clear statement precisely about them: they should not be used in critical GMP applications, and the evidence sections apply only to static models with deterministic output. Outside critical applications, their use remains possible - always with human review of the outputs; responsibility remains with the regulated company. This very boundary is the central point of contention in the consultation - the final text may look different here.

No. Annex 22 is designed as supplementary guidance to Annex 11: the computerized system remains validated under Annex 11; Annex 22 additionally governs the evidence for the embedded model - from intended use to operation.

Annex 22 focuses on the model - training, testing and the evidence of its fitness. We take a wider view: an AI system or AI subsystem integrates one or more models and further functions, such as controls (often called "guardrails"), into a system component. Performance can then be measured at several levels - verifying the fitness of all subcomponents and of their interplay.

Everything that matters: inventory AI applications, document intended use and criticality, define test metrics, secure independent test data, establish monitoring and human oversight. Each of these building blocks follows from Annex 11; ICH Q9(R1) guides the depth of evidence - the draft merely makes them explicit for AI models.

More on the topic

Annex 22 needs the full context.

Annex 22 readiness starts before the final text.

In an initial consultation, we explore what you are planning with AI, where your applications stand today and what Annex 22 means for you. Free of charge, about 30 minutes.

Book an initial consultation