QFINITY updates - publications and committee work
QFINITY · Updates

The latest from QFINITY and the field.

An overview of updates beyond events and projects: articles and publications, work in industry bodies and committees, awards, and our perspective on new regulations.

At a glance

Publications, committee work and more.

Building Trust in AI for Computerized System Validation - QFINITY

Frank Henrichmann, Senior Executive Consultant at QFINITY, spoke with Life Science Connect about the use of AI in computerized system validation. The interview appeared on September 4, 2026, published simultaneously on Pharmaceutical Online, Bioprocess Online and Biosimilar Development.

In the interview, he sets out which validation tasks can be handled reliably by machine. These include checking completeness, tracing requirements to evidence and comparing documents for consistency. It becomes more delicate where someone has to judge whether a piece of evidence is adequate or how a deviation should be assessed. Those judgments stay with people.

He describes two effects as the real danger: automation bias and cognitive ease. Anyone reading plausibly worded suggestions reviews them less rigorously, and that is precisely why putting a person into the process is not enough. It takes independent technical controls that hold even when human review slips. As the methodological framework, he points to the GAMP 5 Second Edition.

Jon O’Connell of Life Science Connect conducted the interview. You can read it here.

Who pushes back when the system speaks? Human Oversight in QA - QFINITY

A follow-up to the 47th GAMP D-A-CH Forum, Berlin, 11 March 2026.

At the 47th GAMP D-A-CH Forum in Berlin in March 2026, Daniel Köpke and Oliver Herrmann asked what Human Oversight means when systems become more intelligent, more complex and more convincing. For us, the answer starts with responsibility: the greatest danger to QA is the silent erosion of visible responsibility, until no one pushes back anymore. This summer, a security incident at the AI provider OpenAI showed how real this question has become. It is the occasion to share the March position now: not foresight, an illustration. This article lays out the position from the talk, and it names what organizations can build to counter it.

Why is this follow-up coming now?

On 26 August 2026, OpenAI published the technical report on an incident from July. In one of the company’s own security tests, agents had broken out of their test environment and compromised real Hugging Face infrastructure. A key factor was that the usual safeguards had been switched off in the test environment for testing purposes. In production, by contrast, they would be in place. Even so, the incident remains a vivid experiment that shows how far the capabilities of these systems now reach. Whether the agents broke out or were let out, both are a failure of architecture, not of a single control. Controls bolted onto a system after the fact cannot keep pace with these systems. Controllability must be designed in before the system speaks. The incident happened far outside the GxP world. The important message here: these AI capabilities are increasingly being used in GxP system landscapes as well. Anyone assessing the associated risks needs to understand both the limits and the potential of these systems.

The starting point is responsibility

Daniel Köpke and Oliver Herrmann presenting at the 47th GAMP D-A-CH Forum in Berlin
BERLINDaniel Köpke and Oliver Herrmann at the 47th GAMP D-A-CH Forum, 11 March 2026.

Every human being has a face and a voice. Both make us recognizable, and both make us responsible: as QA professionals and as people who contribute to patient safety. A patient does not know our systems. They know no SOPs, no validation plans, no model architecture. They trust that in the end a human stands behind quality.

The connection to AI lies in the transition from data to decisions. An AI-enabled system can analyze data, detect patterns, generate test cases, classify deviations and prepare decisions. And as it does, it sounds plausible, often even very plausible.

Everything documented, everything compliant, and no one understands why

In Berlin we opened with a scene. One system recommends release. A second has checked the data. A third has classified the anomaly as acceptable. Everything documented, everything traceable, everything compliant. And no human has really understood why.

This scene is not the future; it is pieced together from today’s everyday practice. A system generates 847 test cases, and the tester checks the output. But who checks the logic behind them, and who decides what was not tested? A chatbot classifies a deviation as minor and proposes the CAPA along with it. The QA professional confirms, and at some point confirming becomes a habit. An autonomous monitoring system reports no trend. But what about the trend outside the pattern the model knows? Silence is not a statement. Silence is an assumption.

Plausibility is not evidence of review

That is the central risk of any Human-in-the-Loop setup. Plausible outputs discourage the thorough review that technical correctness, regulatory robustness and GxP responsibility actually demand. Plausibility can trigger a review. It cannot replace one, and it does not make responsibility transferable. Plausible means: it could be right. Right means: we have checked and understood it. That is not a small difference, it is the difference.

A system bears no regulatory responsibility. It does not know the GxP context the way a human does, it does not recognize when a seemingly correct result becomes dangerous in a critical process, and it does not judge under ambiguity. All of that stays with people. Technical measures such as guardrails can support people in exercising this responsibility. Anyone who trusts these measures blindly, however, perpetuates the very design weakness that makes the Human-in-the-Loop setup vulnerable.

The silent erosion of visible responsibility

The danger comes quietly. Click by click, confirmation by confirmation, buried under ever more decisions waiting to be taken, until no one really pushes back anymore. No single step stands out, and what remains is a QA function that has formally documented everything and in practice no longer decides anything.

When the system speaks, the decisive question remains: who pushes back?

Three sets of rules, one direction

In Berlin we mirrored this question against sets of rules that emerged independently of one another and carry the same expectation. In Article 14, the EU AI Act describes what Human Oversight means for high-risk systems. People must be able to understand, monitor and correct the system, not merely have signed off formally. EU GMP Annex 11 has always demanded controllable, traceable and inspectable computerized systems. It was written before generative AI, and it applies regardless. The draft EU GMP Annex 22 proposes the first formal framework for AI in the GxP environment, with intended use, performance monitoring and change control. In critical applications it currently envisages only static models with deterministic output. The regulated user must hold and review the evidence itself. That applies regardless of whether the model was developed in-house or created with a service provider. On the industry side, GAMP 5 describes the consensus on how these expectations can be implemented. All point in the same direction. Control stays with people.

The role of QA is shifting

Oliver Herrmann presenting at the 47th GAMP D-A-CH Forum in Berlin
ON SITEOliver Herrmann during the talk in Berlin.

A QA function that wants to answer this question shapes the conditions under which human judgment remains effective. In GxP terms, Human Oversight belongs in the intended use, in the business process and in the risk-based controls, with lifecycle evidence that demonstrates its effectiveness. Human Oversight is a capability that is designed into the architecture of the system, not added later in a review step.

In Berlin we described this role in four images. As architecture designer, QA sits at the table when system boundaries are drawn and helps decide which decisions a system may take autonomously and where a human must be able to intervene. As oversight architect, it defines the control points itself instead of leaving them to IT or the vendor: where monitoring takes place, what counts as a deviation, when a system is paused. As escalation designer, it devises the detection paths for silent failures, because drift is not a crash; it creeps. And as guardian of transparency, it records that a validation was incomplete if no one in an audit can explain why the system design was chosen, why human oversight was defined as it was and why the decision in operation was made as it was.

The sentence the talk was building toward still stands. Future-proof QA does more than validate systems. It validates that people remain capable of deciding.

And it needs people who can review. Four conditions determine whether pushback happens in daily work:

1
Competence. Whoever reviews must understand what the system can and cannot do, know the context in which it is used and be able to judge its limits.
2
Time. A review with no time set aside for it becomes a confirmation.
3
Psychological safety. Pushing back against a result that sounds plausible and that everyone accepts requires an environment that explicitly expects dissent and makes it possible without repercussions.
4
Real authority. Whoever reviews must be allowed to stop a system, formally and in everyday practice. A system without a defined stop is not a controlled system.

Organizations must design Human Oversight so that responsibility is exercised in daily work and its effectiveness remains demonstrable. And that assignment of responsibility belongs on record. It is not the system that decided. A human took responsibility for the system’s decision, with name, role, qualification and (digital) signature. In an audit there is no line for the model. A responsibility that exists only on paper protects no patient.

Three supporting voices, one shared core

Since the talk, this position has not stood alone. The rapporteur of the Annex 22 drafting group in Barcelona, a National Expert at the FDA in Boston and industry at the EMA expert workshop arrived independently at the same line. The synthesis is in our article Three Signals, One Line. How Human Oversight can be set up and checked is covered in AI Governance and Human Oversight.

An AI strategy that anchors Human Oversight and visible responsibility is a good starting point. More important still is the attitude: the system works for us, not the other way around.

For us, that is the core of Digital Compliance: quality is not merely demonstrated; it is designed to be trustworthy and to grow with knowledge. Trust has a face and a voice. Our job is to make sure neither disappears into our systems.

Further reading: Rückblick: GAMP D-A-CH in Berlin (ISPE D-A-CH, 11 March 2026, in German)OpenAI: report on the Hugging Face security incident (26 August 2026)

Three Signals, One Line: AI in the GxP Environment from Barcelona and Boston to the EMA Workshop - QFINITY

Within seven months, three signals converged on how to assess AI in the GxP environment: the rapporteur of the EMA drafting group for EU GMP Annex 22 explained the draft’s criticality logic in Barcelona in December 2025, a National Expert at the US FDA made clear in Boston in June 2026 that the rules still hold, and at the EMA expert workshop on 30 June 2026 industry answered with a joint position. QFINITY followed all three, Barcelona and Boston on site, the EMA workshop via its public broadcast. The occasions were independent of one another, yet all of them arrive at the same five sentences. It is the line we ourselves presented at the GAMP D-A-CH Forum in Berlin in March 2026, with the question: who pushes back when the system speaks?

Taken in turn: in Barcelona the drafting group’s rapporteur spoke, in Boston a National Expert at the FDA, at the workshop industry addressed the EMA. At the end we put the three signals in perspective.

Barcelona, December 2025: How the drafting group thinks about criticality

At the 2025 ISPE Pharma 4.0 Conference (9 and 10 December 2025, Barcelona), the rapporteur of the Annex 22 drafting group, a representative of the Danish Medicines Agency, explained how the draft delimits its scope. The guiding question was: what effect would an error have, and would it be detected? Which technology is in use makes no difference to that classification, at least at first. An AI-supported application whose output passes through an expert review anyway, for example training material or SOP drafts, counts as non-critical. An application whose output feeds into the quality decision without further review, for example in quality control or automated visual inspection, counts as critical.

The rapporteur then explained the draft’s original regulatory intent. Within the critical area, the draft initially excludes certain technologies. On the slide this area sat in the top right, and as the “upper right-hand corner” it became a catchphrase among experts. Dynamic systems that keep learning in operation are left out, as are probabilistic systems where the same input and the same version do not guarantee the same result. Consequently, that also applies to large language models. Although this view starts from process and system design, it was in the end tied to technology. That became one of the main points of discussion across the industry. He had delivered the same message with the same slides in the GAMP D-A-CH community a few days earlier: on 4 December 2025 at the 2nd GAMP Conference “Künstliche Intelligenz trifft Pharma” in Mannheim. QFINITY was involved in leading the GAMP D-A-CH community for more than a decade and helped build the AI community in D-A-CH.

The second thought from Barcelona concerns evidence. A trained model cannot be proven out by a single deterministic test. The evidence takes a different form: test data that are themselves subject to requirements, and metrics that carry the evidence for control and effectiveness. That is how data scientists think, and it is at the same time the principle of quality risk management: decision under uncertainty. On evidence, then, Annex 22 imports no foreign logic into the GxP world; it applies the existing logic to models. On scope, by contrast, the draft draws the line a priori rather than judging a model type’s admissibility on the basis of the risk assessment. Industry would later take the discussion up from there.

Boston, June 2026: An FDA voice says the rules still apply

At the ISPE AI in Life Sciences Summit in Boston (22 and 23 June 2026), Seneca Toms, National Expert for Drugs at the US FDA, spoke about how industry is handling AI. Oliver Herrmann was in the room. Frank Henrichmann, as Chair of the GAMP Global Steering Committee, represented the “Powered by GAMP” side of the summit. What made the talk convincing was the clarity with which a regulator’s voice applied the old principles to the new technology. Safe and effective products, controlled processes, identified and managed risks, scientifically justified decisions: that held before AI, and it holds after. ISPE’s editorial team summarized the talk in an August iSpeak post. It notes explicitly that the summary has not been vetted by any of the agencies mentioned and does not represent an official agency position. We therefore present the thoughts that follow as a reflection on the talk, not as an FDA statement.

We pick up four thoughts from Boston because they apply directly to regulated companies. How deeply you test follows the decision a system supports: a tool that summarizes meeting notes needs a different level of assurance than a system that feeds into decisions on product quality or patient safety. Oversight begins with understanding; whoever approves a result without understanding it is not exercising Human Oversight. The greatest risk sits in trust. Toms described inspections where systems had not failed; people had simply stopped asking, because the systems had been running for years. It reflects an observation we also described in Berlin. We will come back to it below. With AI the pattern repeats as soon as recommendations are accepted because they are convenient or look credible. And the “current” in cGMP demands keeping pace. New tools are measured against today’s state, because paper, legacy systems, people and today’s means of controlling AI all have limits.

“You can outsource a lot of things, but you cannot outsource your common sense.” (Seneca Toms, US FDA, as quoted by ISPE iSpeak, 24 August 2026)

EMA expert workshop, 30 June 2026: Industry answers with one voice

One week after Boston, the EMA spent a day listening to industry. At the expert workshop on the draft EU GMP Annex 22, experts nominated by the associations presented their positions on six topics set by the EMA, the six pillars of the EMA’s guardrail architecture, from regulatory pathways for adaptive models through Human Oversight, validation and lifecycle to cybersecurity. We followed the publicly broadcast first day in full. The workshop followed the 2025 consultation, which drew 1,359 comments from 79 organizations; the call for a risk-based approach was its clearest theme. On the second, non-public day the drafting group took the input on board and continued its work on the text. The two-day sequence was set from the outset. For us, the signal lies in the tone of the public statement the EMA gave afterwards. It suggests that the ideas and concepts presented were received as helpful. A senior FDA official, too, publicly praised the workshop’s format and dialogue.

The associations had been asked to present divergent views as well. The outcome was nonetheless clear: their approaches agree, and they come down to how a quality-risk-based approach is interpreted. On the central question of scope, industry’s position departed from the draft. The draft excludes dynamic, probabilistic and generative models from critical applications. Industry countered that no model type is inadmissible or harmless per se. Admissibility is decided by the risk assessment in the specific use case. On Human Oversight, industry proposed replacing the Human-in-the-Loop mechanism fixed in the draft with a Human Oversight concept with several forms. The range runs from approval of every output to ongoing monitoring with intervention by exception. Which form is appropriate follows from the risk assessment. And on guardrails, industry drew the line itself: they reduce risk, they do not remove it, and a control that is meant to lower risk needs its own evidence of effectiveness.

From QFINITY’s point of view, the quiet highlight was an architecture diagram shown at the workshop: the AI subsystem of model, integration code and guardrails as a part inside the computerized system, which also includes process and people. That embedding is precisely the architecture behind our validation of AI in the GxP environment: the model is verified; the computerized system as a whole is validated in the process.

Five sentences all three share

Placed side by side, the three occasions leave a common core that none of the voices disputes:

1
Criticality is derived from the process and measured by impact and detectability, not by technology. Whether framed as impact and detectability, as the significance of the decision or as the risk assessment in the use case, all three describe the same axis. The particular traits of generative or dynamic models belong one level down, in the functional risk assessment, where guardrails come in as controls in the sense of quality risk management.
2
Human Oversight is a capability. Barcelona makes expert review the measure of criticality, the FDA voice from Boston demands understanding rather than mere approval, industry proposes replacing the fixed HITL mechanism with a Human Oversight concept with several forms, and all three presuppose that people can review effectively.
3
The form of evidence shifts to statistics, and it stays within GxP logic. A model is verified with test data that carry their own requirements, with metrics and with confidence levels, and that follows the principle of decision under uncertainty.
4
Oversight applies to the whole lifecycle. It runs from planning and design through initial verification and the whole period of use to decommissioning. That includes adjusting the form of oversight, and it includes monitoring. Confidence in a system is not established once and then left alone.
5
Accountability stays with the operating company. No model and no service provider relieves you of it. Toms said it from the FDA’s perspective, industry presented it as consensus at the workshop, and your next inspection will assume it.

Our position from Berlin: Who pushes back when the system speaks?

The five sentences match the position Daniel Köpke and Oliver Herrmann presented at the 47th GAMP D-A-CH Forum in Berlin on 11 March 2026. They asked what Human Oversight means when systems become more intelligent, more complex and more convincing. The starting point was responsibility: a patient knows neither SOPs nor validation plans; they trust that in the end a human stands behind quality. An AI-enabled system can analyze data, detect patterns, classify deviations and prepare decisions, and it sounds plausible while doing so. That is exactly where the risk lies: plausibility is not evidence of review. It can trigger a review, it cannot replace one, and it does not make responsibility transferable.

The biggest danger to QA is therefore not AI. It is the silent erosion of visible responsibility: click by click, confirmation by confirmation, until no one pushes back anymore. The role of QA shifts accordingly: it does not just validate systems, it shapes the conditions under which human judgment remains effective. Human Oversight belongs embedded in intended use, business process, data integrity, risk-based controls and lifecycle evidence, so that responsibility is not merely documented but exercised, and its effectiveness stays demonstrable. The full version of this position is in Who Pushes Back When the System Speaks?

What conditions do you need to create today so that in three years someone will still challenge a recommendation the system has delivered without complaint all along?

What follows for regulated companies

The convergence has a practical side. Anyone working by these five sentences today is unlikely to have to rebuild for the final Annex 22, whether the EMA follows industry’s risk principle or keeps the exclusion of certain model classes. The draft’s evidence logic, from intended use through independent test data to monitoring, holds in every outcome of the revision. And it holds before an FDA inspection too, because what counts there is what Toms named in Boston: understanding, risk, lifecycle, accountability. That evidence logic is just as necessary for systems to deliver the expected performance and, with or without an AI label, make a tangible contribution to relief and value.

The entry point is the criticality question: which AI-supported applications deliver results that feed without further review into quality decisions that touch patient safety, product quality or data integrity? The effectiveness of Human Oversight follows from there. What matters is less whether a review step is documented than whether the reviewing person understands the context of use, knows the system’s limits and is free to disagree. How that can be checked is described under AI Governance and Human Oversight. This includes the question of whether dissent still occurs in day-to-day operation.

What comes next

For Annex 22, a workshop report has been announced first; a revised draft is expected afterwards. Our Annex 22 page sets out in three questions what the final text will turn on; as soon as the report is available, we will measure it against them. Until then, the position is a plain one: the computerized system is validated under Annex 11, quality risk management guides how deep the evidence needs to go, and Toms describes no different expectation from inspections. For QFINITY, the three signals from regulators and industry confirm the path we presented in Berlin: AI continues the line of CSV and CSA. That is how we described it in Pharmaceutical Engineering in January, and that is how we read this year’s regulator and industry signals.

Further reading: US FDA on AI, Critical Thinking, and the Enduring Principles of Quality (ISPE iSpeak, 24 August 2026)Human-in-the-Loop as an Illusion of Control? (Herrmann and Henrichmann, ISPE iSpeak, 4 September 2026)Chapter 4, Annex 11, Annex 22: Three Drafts, One Control System (QFINITY)