EMA Workshop Report on Annex 22: The Drafting Group Has Discussed Opening the Scope to Generative AI

EMA-Workshop-Report zu Annex 22: sechs Ströme werden zu vier Linien und laufen zu einer Aussage zusammen (QFINITY)

In early October 2026 the EMA published the report on its expert workshop of 30 June 2026. In 15 pages, document EMA/156789/2026 records what the industry associations presented on six topics, and it contains one sentence that reaches beyond the workshop: the drafting group has discussed widening the scope of EU GMP Annex 22 to dynamic, adaptive, probabilistic and generative models. The condition would be a "fully documented, robust, risk-based control strategy". Here is our reading of what the report confirms and what it leaves open.

We followed the publicly broadcast workshop day and brought the signals from Barcelona, Boston and the workshop together in our August article Three Signals, One Line. With the report, a written account of the workshop by the EMA itself is available for the first time.

The sentence that matters

The 2025 consultation draft excludes generative AI and Large Language Models from its scope and, like dynamic and probabilistic models, does not provide for them in critical GMP applications. The report quotes that sentence in its introduction and sets the consultation feedback against it. The comments supported allowing such models in GMP applications, "whether critical or non-critical". The drafting group, the EMA writes, has therefore discussed widening the scope, "provided they comply with the requirements of the annex and are supported by a fully documented, robust, risk-based control strategy".

Two things belong to that sentence. First, it records a discussion, not a decision. As the next step, the report names the revision of the draft by the drafting group, without giving a date. Second, the question of which elements such a control strategy needs was the very reason for the workshop. The drafting group wanted to hear from experts whether a risk-based approach can be applied to generative AI and which control mechanisms would carry it.

Four lines across six topics

For each topic, the report gives the associations' positions and the regulators' questions. Each topic closes with key messages and a section on the potential impact on the annex text. Four lines run through the topics.

1
No prohibition by technology category. The consolidated industry position opposed categorical exclusions. Whether a model is acceptable should follow from intended use, decision consequence, model influence, uncertainty and complexity, and from whether the residual risk can be reduced to an acceptable level. The report records as a key message that existing quality risk management principles apply to adaptive and probabilistic models as well. A valid outcome of that assessment may still be not to use the model. That conclusion should be documented.
2
Human oversight rather than human-in-the-loop by default. The report distinguishes oversight as a lifecycle-wide objective from human-in-the-loop as one possible implementation, in which a process waits for a human action. Its key messages state that human-in-the-loop is not a default requirement for all AI uses, that oversight must be evidenced and periodically reviewed, and that it cannot compensate for inadequate validation. The regulators asked how automation bias can be managed and how the performance of the reviewing person can be monitored.
3
Control mechanisms need evidence. The report documents the presented layering of prevention at the input, detection in the model and containment at the output. It also notes that some of the five use cases were pilots or theoretical frameworks. Two points from the discussion are recorded. Controls designed for known failure modes do not by themselves cover unanticipated ones, and in the annex itself industry would prefer the term control mechanism over guardrail, because guardrails are technology-specific and transient. On that position, the annex should name objectives such as prevention, detection and containment, while technical methods belong in an updateable format.
4
Responsibility stays with the regulated company. That covers the use of the model as much as the management of suppliers and cloud providers. On the industry position reported, AI systems consist of data, model and hardware and therefore remain computerized systems. Annex 11 applies to them. The report refers to its chapter on outsourced activities; in the current Annex 11 the subject sits in section 3, in the 2025 draft in section 7. Whoever uses a third-party model through an interface needs access to the evidence with which its behavior can be measured. If the provider does not supply it, the use cannot be justified.

What has not been worked out yet

What the workshop left open is just as telling. Some of the use cases presented were, according to the report, not yet production systems but pilots or concepts. On the question of how independent test data can be preserved for a model that keeps learning in operation, the report records that the answer did not define one mandatory technical approach and that implementation for continuously learning systems needs clarification. And for the particular behaviors of generative models, hallucination, fabrication, overconfidence in the output, the workshop set no method for estimating rates or confidence. The revised draft will have to settle these questions, or hand them openly to the operator's risk assessment.

Our reading

For QFINITY, the report overlaps with the architecture we describe on our Annex 22 page and in Validation of AI in the GxP environment. The model is a component with its own evidence, and it is verified. The computerized system in which it works with control mechanisms, process and people is validated in its process. The report does not commit to this distinction. It uses "validation" for lifecycle and system questions as well and keeps the objects of evidence open side by side: in its outlook on the validation topic it writes "revalidation/re-verification" and names an "AI validation/qualification package" as a possible consequence, mirroring the industry presentation rather than stating a regulatory position. The conceptual core of what an operator should be able to show is nevertheless visible in the report: evidence of model performance, information on the training data, controls around input and output, and evidence that the system works as expected in operation. In substance that matches the evidence we have described since the GAMP workshop at the ISPE summit in Boston in June, on whose core team Frank Henrichmann worked for QFINITY: intended use in the company's own process, acceptance criteria, a test set kept separate from training, the supplier's model card, and the residual risk the operator carries.

On oversight, the report meets the line we presented in March at the GAMP D-A-CH Forum under the question Who Pushes Back When the System Speaks? and deepened in September for ISPE iSpeak under the title Human-in-the-Loop as an Illusion of Control? Human-in-the-loop alone does not establish control. Control comes about when the reviewing person understands the context of use, knows the limits of the model and is allowed to push back, and when that capability can be evidenced in operation. In the report, one speaker puts it this way: oversight provides knowledge, detection and triggers for action, and it is the action that reduces the risk.

Which of your AI-supported applications would you operate today on the basis of a documented, risk-based control strategy? And for which would the assessment conclude not to use them?

A checklist for operators

The revised draft has no date. Until then the 2025 consultation draft remains the reference for preparation, not an annex in force, and the report does not change its wording. Anyone planning an AI application in a GMP environment today can use the four lines of the report as a checklist. It starts with a criticality assessment that, beyond direct impact, takes in decision consequence, model influence, uncertainty and the detectability of errors. The form of oversight must fit the risk assessment, and its effectiveness must be demonstrable. The evidence for the control mechanisms rests on the overall package of controls rather than on each mechanism in isolation. Supplier agreements secure access to evidence, participation in change control and transparency of the controls. How this can be anchored in AI governance with human oversight is described on our page on the subject. We have also worked the state of the report into the three open questions on our Annex 22 page.

Entry offer
Readiness Assessment and Roadmap: Chapter 4, Annex 11 & 22

In the AI module we assess your AI applications along the four lines of the report: criticality, form of oversight, evidence for control mechanisms, supplier agreements.

Modules
Data, systems, AI
Duration
Three to six weeks
Result
Control map, gap list, roadmap

View the offer →

Further reading: EMA: Report of the multistakeholder workshop on expert contributions to AI guidance development (Annex 22), EMA/156789/2026↗EMA: workshop page with agenda and report↗Three Signals, One Line: AI in the GxP environment between Barcelona, Boston and the EMA workshop (QFINITY)↗