Who Pushes Back When the System Speaks?
A follow-up to the 47th GAMP D-A-CH Forum, Berlin, 11 March 2026.
At the 47th GAMP D-A-CH Forum in Berlin in March 2026, Daniel Köpke and Oliver Herrmann asked what Human Oversight means when systems become more intelligent, more complex and more convincing. For us, the answer starts with responsibility: the greatest danger to QA is the silent erosion of visible responsibility, until no one pushes back anymore. This summer, a security incident at the AI provider OpenAI showed how real this question has become. It is the occasion to share the March position now: not foresight, an illustration. This article lays out the position from the talk, and it names what organizations can build to counter it.
Why is this follow-up coming now?
On 26 August 2026, OpenAI published the technical report on an incident from July. In one of the company’s own security tests, agents had broken out of their test environment and compromised real Hugging Face infrastructure. A key factor was that the usual safeguards had been switched off in the test environment for testing purposes. In production, by contrast, they would be in place. Even so, the incident remains a vivid experiment that shows how far the capabilities of these systems now reach. Whether the agents broke out or were let out, both are a failure of architecture, not of a single control. Controls bolted onto a system after the fact cannot keep pace with these systems. Controllability must be designed in before the system speaks. The incident happened far outside the GxP world. The important message here: these AI capabilities are increasingly being used in GxP system landscapes as well. Anyone assessing the associated risks needs to understand both the limits and the potential of these systems.
The starting point is responsibility

Every human being has a face and a voice. Both make us recognizable, and both make us responsible: as QA professionals and as people who contribute to patient safety. A patient does not know our systems. They know no SOPs, no validation plans, no model architecture. They trust that in the end a human stands behind quality.
The connection to AI lies in the transition from data to decisions. An AI-enabled system can analyze data, detect patterns, generate test cases, classify deviations and prepare decisions. And as it does, it sounds plausible, often even very plausible.
Everything documented, everything compliant, and no one understands why
In Berlin we opened with a scene. One system recommends release. A second has checked the data. A third has classified the anomaly as acceptable. Everything documented, everything traceable, everything compliant. And no human has really understood why.
This scene is not the future; it is pieced together from today’s everyday practice. A system generates 847 test cases, and the tester checks the output. But who checks the logic behind them, and who decides what was not tested? A chatbot classifies a deviation as minor and proposes the CAPA along with it. The QA professional confirms, and at some point confirming becomes a habit. An autonomous monitoring system reports no trend. But what about the trend outside the pattern the model knows? Silence is not a statement. Silence is an assumption.
Plausibility is not evidence of review
That is the central risk of any Human-in-the-Loop setup. Plausible outputs discourage the thorough review that technical correctness, regulatory robustness and GxP responsibility actually demand. Plausibility can trigger a review. It cannot replace one, and it does not make responsibility transferable. Plausible means: it could be right. Right means: we have checked and understood it. That is not a small difference, it is the difference.
A system bears no regulatory responsibility. It does not know the GxP context the way a human does, it does not recognize when a seemingly correct result becomes dangerous in a critical process, and it does not judge under ambiguity. All of that stays with people. Technical measures such as guardrails can support people in exercising this responsibility. Anyone who trusts these measures blindly, however, perpetuates the very design weakness that makes the Human-in-the-Loop setup vulnerable.
The silent erosion of visible responsibility
The danger comes quietly. Click by click, confirmation by confirmation, buried under ever more decisions waiting to be taken, until no one really pushes back anymore. No single step stands out, and what remains is a QA function that has formally documented everything and in practice no longer decides anything.
When the system speaks, the decisive question remains: who pushes back?
Three sets of rules, one direction
In Berlin we mirrored this question against sets of rules that emerged independently of one another and carry the same expectation. In Article 14, the EU AI Act describes what Human Oversight means for high-risk systems. People must be able to understand, monitor and correct the system, not merely have signed off formally. EU GMP Annex 11 has always demanded controllable, traceable and inspectable computerized systems. It was written before generative AI, and it applies regardless. The draft EU GMP Annex 22 proposes the first formal framework for AI in the GxP environment, with intended use, performance monitoring and change control. In critical applications it currently envisages only static models with deterministic output. The regulated user must hold and review the evidence itself. That applies regardless of whether the model was developed in-house or created with a service provider. On the industry side, GAMP 5 describes the consensus on how these expectations can be implemented. All point in the same direction. Control stays with people.
The role of QA is shifting

A QA function that wants to answer this question shapes the conditions under which human judgment remains effective. In GxP terms, Human Oversight belongs in the intended use, in the business process and in the risk-based controls, with lifecycle evidence that demonstrates its effectiveness. Human Oversight is a capability that is designed into the architecture of the system, not added later in a review step.
In Berlin we described this role in four images. As architecture designer, QA sits at the table when system boundaries are drawn and helps decide which decisions a system may take autonomously and where a human must be able to intervene. As oversight architect, it defines the control points itself instead of leaving them to IT or the vendor: where monitoring takes place, what counts as a deviation, when a system is paused. As escalation designer, it devises the detection paths for silent failures, because drift is not a crash; it creeps. And as guardian of transparency, it records that a validation was incomplete if no one in an audit can explain why the system design was chosen, why human oversight was defined as it was and why the decision in operation was made as it was.
The sentence the talk was building toward still stands. Future-proof QA does more than validate systems. It validates that people remain capable of deciding.
And it needs people who can review. Four conditions determine whether pushback happens in daily work:
Organizations must design Human Oversight so that responsibility is exercised in daily work and its effectiveness remains demonstrable. And that assignment of responsibility belongs on record. It is not the system that decided. A human took responsibility for the system’s decision, with name, role, qualification and (digital) signature. In an audit there is no line for the model. A responsibility that exists only on paper protects no patient.
Three supporting voices, one shared core
Since the talk, this position has not stood alone. The rapporteur of the Annex 22 drafting group in Barcelona, a National Expert at the FDA in Boston and industry at the EMA expert workshop arrived independently at the same line. The synthesis is in our article Three Signals, One Line. How Human Oversight can be set up and checked is covered in AI Governance and Human Oversight.
An AI strategy that anchors Human Oversight and visible responsibility is a good starting point. More important still is the attitude: the system works for us, not the other way around.
For us, that is the core of Digital Compliance: quality is not merely demonstrated; it is designed to be trustworthy and to grow with knowledge. Trust has a face and a voice. Our job is to make sure neither disappears into our systems.
Further reading: Rückblick: GAMP D-A-CH in Berlin (ISPE D-A-CH, 11 March 2026, in German)↗OpenAI: report on the Hugging Face security incident (26 August 2026)↗




