A reviewer facing a wall of AI-generated results - Human-in-the-Loop as a control - QFINITY
QFINITY · ISPE iSpeak

A human in the loop is not yet a control.

In their article for ISPE iSpeak of 4 September 2026, Frank Henrichmann and Oliver Herrmann (QFINITY) show why Human-in-the-Loop in AI-supported GxP processes often offers less protection than is assumed in practice: anyone who only reviews results instead of producing them moves into a purely supervisory role, in which errors are noticed less often. From this, the article derives a multi-layered control concept.

The article

Human-in-the-Loop comes with preconditions.

Human-in-the-Loop (HITL) is a risk-based control concept in which qualified people review, approve or override the results of automated or AI-supported GxP processes. Drawing on the research on human factors and automation bias, Henrichmann and Herrmann test whether HITL is inherently a reliable control. The conclusion is sober: as soon as the human moves from active analysis into a purely supervisory role, the likelihood of catching errors declines. Whether the control works is decided by its design.

Detectability does not come from someone looking. It comes from the process and the system enforcing it.
What the article establishes

What hollows out Human-in-the-Loop.

  • Automation bias

    Recommendations from the system carry more weight than one's own judgment, even when contradictory information is available. Language models that uncritically confirm what their users say can amplify the effect, because an affirmative tone dampens the motivation to question the result.

  • Plausibility trap

    Fluent, formally consistent results create cognitive ease. What is easy to process feels plausible. A high review volume amplifies this effect, because reviewers fall back on sampling and a superficial plausibility check.

  • Review fatigue

    Anyone reviewing large volumes of uniform, mostly correct results loses vigilance. In medicine, the related phenomenon of automation complacency is discussed in the context of decision-support systems; AI-supported GxP workflows create structurally comparable conditions.

  • Attitude effects

    A person's attitude toward automation is a strong predictor of whether errors are detected and corrected. Formal qualification alone does not ensure the effectiveness of the control as long as attitude and working conditions remain unaddressed.

  • Trust calibration

    Excessive distrust is harmful too. Anyone who distrusts the system and intervenes unnecessarily discards correct results and introduces new errors. The goal is a level of trust that matches the system's demonstrated performance.

  • Human-AI loop and silent delegation

    Practitioner literature describes, so far without robust studies, how documentation and decision practices can adapt to the logic of the system. In parallel, organizations tacitly treat the system as the party that carries responsibility while formal accountabilities remain unchanged.

CSV use cases

Where the psychology becomes operational.

Two examples from computerized system validation show how the effects would play out.

AI-supported test review

An AI-enabled tool rates an OQ result as passed because the acceptance criteria are formally met. The result looks correct, and the reviewer approves it. Nobody has checked the unspoken assumption that the dataset represents the production configuration. An outdated configuration state is technically a simple error and still hard for the reviewer to see. The risk grows when one language model generates test cases and a second one reviews them, because both can follow similar patterns. Tests from language models also often degrade into smoke tests that confirm the flow without verifying the requirement behind it.

Traceability agents

An agent that cross-checks requirements, specifications and test artifacts reports the traceability of an audit trail requirement as complete: a test case references the requirement, and the test mentions the function. The green status says nothing about whether completeness, immutability and attribution of the audit trail were verified at all. The clean artifact still counts as evidence.

The obvious reaction, adding further review stages, does not solve the problem. In controlled experiments, correction rates fell as soon as corrections took effort. Human review then risks becoming a ritual, in the words of the article formally compliant but substantively hollow. As the system matures, the visible error rate may also fall while the review volume grows. That need not mean the risk is decreasing; it may mean that errors have become harder to detect. Which reading is true only becomes clear by looking at model quality, the review process and the control design.

Control system

Detectability is designed in, not assumed.

The regulatory direction is unambiguous. The AI Act, Regulation (EU) 2024/1689, requires effective human oversight for high-risk AI systems, while EU GMP Annex 11 requires clearly defined responsibilities and risk-based controls. The draft Annex 22 identifies human oversight as a key control mechanism. An April 2026 FDA warning letter to Purolea Cosmetics Lab cites the use of AI agents that generated GMP records without sufficient verification. The authors answer with a layered model. HITL remains one layer in a control system of people, processes and technology, and the technical controls of that control system keep working even when vigilance declines. How vendors and regulated users divide these requirements in practice is the subject of our recap of the SAP expert workshop.

Risk driverTypical failure of the controlWhat compensates for it (examples after Henrichmann/Herrmann, ISPE iSpeak 2026)
Automation biasProposals from the system are adopted without anyone challenging themMandatory own assessment and justification, review by a second independent party, guardrails in the system
Plausibility trap (cognitive ease)The reviewer confirms what looks plausible instead of analyzing itAcceptance and rejection must be justified; the interface shows uncertainty instead of traffic lights
Review fatigueApprovals become routine at high throughputLimited review volumes per person, rotation of reviewers, risk-based triggers instead of full review
Human-AI loopDocumentation and decisions imperceptibly align with the logic of the systemMonitoring of drift and data quality, periodic re-verification of the training data
Attitude effectsA person's attitude toward automation shapes review behaviorBuilding AI literacy, peer review, changing reviewers

Behind this sits the risk logic of ICH Q9(R1): severity, probability, detectability. Human-in-the-Loop only ever claims the third factor, and that is exactly the one that has to be demonstrated. Guardrails belong in the operational frame of the system, meaning system prompts, orchestration, tool rules and execution constraints. Their execution and every exception are logged. That keeps them observable and open to review, and their effectiveness is demonstrated across the lifecycle. Only then do guardrails ease the burden on human review, following the same logic as transferring a human task to automation: the control moves, the evidence stays. That evidence does not end at release. In operation it continues through ongoing process monitoring with drift and data quality control. For results that can be objectively derived from source data and reproduced, human review can be lighter; where the system delivers conclusions, it becomes more intensive. Classifying a result as objectively derivable does not replace the review: results flagged that way can trigger a particularly strong automation bias.

Regulation (EU) 2024/1689 EU GMP Annex 11 Annex 22 (draft) ISPE GAMP AI Guide
The authors

From the practice of the GAMP committees.

Frank Henrichmann, Senior Executive Consultant at QFINITY, is Chair of the GAMP Global Steering Committee and a member of the global GAMP AI SIG. Oliver Herrmann, founder and CEO of QFINITY, is Chair of the ISPE GAMP Europe Steering Committee and a member of the ISPE AI Community of Practice. The article picks up the argument both began in Pharmaceutical Engineering in January/February 2026: AI as the next stage of validation, with the expert who remains indispensable in the loop.

GAMP Global Steering Committee GAMP AI SIG ISPE GAMP Europe Steering Committee ISPE AI Community of Practice
Read the full article at ISPE iSpeak
Related topics

From oversight to demonstrable control.

Is your Human-in-the-Loop a control or a gesture?

Bring one AI-supported process, such as test review or traceability. In the initial consultation we map it to the mechanisms and name the technical control that is missing once attention fades. Free of charge, about 30 minutes, directly with one of the authors.

Book an initial consultation