Human Oversight Is a Control You Qualify, Not a Sentence in Your Risk Assessment
Every AI risk assessment in pharma contains the sentence "a qualified human reviews the output". Draft Annex 22 turns that sentence into a monitored manual process with its own training record, its own performance data and its own way of lapsing.
Open any pharmaceutical AI risk assessment written in the last two years and you will find a load-bearing sentence: a qualified reviewer checks the output before it is used. That sentence usually does the work of an entire validation argument. It is why the test set is smaller than a data scientist would like, and why the system was classified as decision support rather than decision maker.
The draft that Europe put on the table in July 2025 reads that sentence back to you and asks what you did about it. Clause 3.3 of draft Annex 22 says that where the effort to test a model has been diminished because a human operator is in the loop, "the training and consistent performance of the operator should be monitored like any other manual process". Clause 10.5 adds that records must be kept of that review, and that depending on criticality, this "may imply a consistent review and/or test of every output from the model, according to a procedure".
That is not a governance nicety. It is a second operating obligation, priced in headcount and telemetry, taken on the moment you used a human to justify testing less. Almost nobody budgets for it.
- Draft Annex 22 clause 3.3 makes the reviewer a monitored manual process when human review was used to reduce model testing. Clause 10.5 says this may extend to review of every output.
- Annex 22 is a draft on 30 August 2026. Consultation ran 7 July to 7 October 2025. The binding EU text remains the 2011 Annex 11.
- The EU AI Act names automation bias in Article 14(4)(b); the 7 July 2025 Annex 22 draft never uses the term. You have to import the concept yourself.
- Qualification is a compound key — person, system, model version, use case. Change any component and the earlier evidence stops covering the new arrangement.
- Two numbers make the control auditable: override rate by output type and median review time. If you cannot produce them, you have a slogan.
What draft Annex 22 actually says about the human in the loop
The two clauses sit in different chapters and are usually read separately. Clause 3.3 is under Intended Use: it requires the intended-use description to include the responsibility of the operator, then imposes the monitoring duty. Clause 10.5 is under Operation and is about records: where testing was diminished and a human is in the loop, "records should be kept from this process".
Read together they describe a bargain. You may spend less on model testing if a competent human is genuinely in the path. In exchange, the human becomes a qualified GMP process step, with training evidence, performance evidence and records, and the regulator reserves the right to expect that every single output was looked at. The full clause-by-clause reading of the draft annex covers the scope test that decides whether any of this applies to your system at all.
Two adjacent clauses tighten it. Clause 4.3, headed "No decrease", says acceptance criteria "should be at least as high as the performance of the process it replaces" — so you need a documented performance figure for the human-only process before you displace it. Clause 10.1 puts the model, the system and "the whole process it is automating or assisting" under change control, with any decision not to retest after a change fully justified.
The industry noticed which of these clauses had teeth. In its consultation response filed on 7 October 2025, ISPE argued that clause 10.5 "contains too much HOW and is considered too prescriptive", warning that "such requirements may discourage the use of ML", and proposed replacing the every-output language with a quality-risk-management formulation. Across more than twenty separate commented line ranges in that submission, ISPE raised nothing at all on lines 52 to 56 — clause 3.3, the one that turns the operator into a monitored manual process. The obligation nobody argued about is the one that costs most to run.
Which instrument binds you on 30 August 2026
This is where internal AI policies most often misstate the position, and it is the fastest thing for an auditor to catch.
| Instrument | Relevant provision | Status on 30 Aug 2026 |
|---|---|---|
| EU GMP Annex 11 (2011) | Computerised systems, risk-based validation | Binding |
| EU GMP Chapter 2 (in operation 16 Feb 2014) | 2.11 — continuing training, effectiveness periodically assessed, records kept | Binding |
| Draft Annex 22 | 3.3 operator monitoring; 10.5 human review records | Draft; consultation closed 7 Oct 2025 |
| Draft Annex 11 revision | Expanded lifecycle, audit trail, access control | Draft; same consultation |
| EU AI Act, Reg. (EU) 2024/1689 | Art. 4 AI literacy; Art. 14 human oversight; Art. 26(2) deployer duty | Art. 4 live since 2 Feb 2025; Art. 14 and 26 apply on the high-risk timetable |
| Reg. (EU) 2026/1744 (Digital Omnibus on AI) | Defers high-risk application dates | In force 27 Jul 2026 |
| 21 CFR 211.22(c) | Quality unit approves procedures and specifications | Binding in the US since 1978 |
The last row of that timetable moved only weeks ago. Regulation (EU) 2026/1744 of 8 July 2026 was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026, six days before the AI Act's original high-risk deadline. It pushes Annex III standalone high-risk systems to 2 December 2027 and Annex I product-embedded systems to 2 August 2028. The deferral is selective: prohibited practices, the Article 4 AI literacy duty and the general-purpose AI obligations are live and untouched.
For most manufacturing and quality AI, Article 14 is not the binding hook, because that work is rarely an Annex III category. It is still the clearest published statement of what a regulator means by human oversight, and inspectors read it. The pillar on what actually binds AI validation in 2026 versus what is still draft lays out the full stack.
Automation bias is the failure mode, and only one of these texts names it
Article 14(4)(b) of the AI Act requires that the person assigned oversight is enabled "to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)". Article 26(2) puts the matching duty on the deployer: "Deployers shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support."
Now the contrast. Searching the 7 July 2025 consultation version of Annex 22 on 30 August 2026, "automation bias" appears zero times across its six pages, and so does "override". The annex invokes the human-in-the-loop arrangement four times without naming the mechanism by which it fails. That concept has to enter your design from outside the annex.
Automation bias is not speculative. In a reader study published in Radiology in 2023, Dratsch and colleagues had 27 radiologists assign BI-RADS categories to 50 mammograms with and without a purported AI suggestion. When the suggestion was wrong, the proportion of correctly rated mammograms fell from 79.7% to 19.8% for inexperienced readers and from 82.3% to 45.5% for the most experienced.
Sit with that second figure. The most experienced group, doing the task they do professionally, was wrong more than half the time when the machine was wrong. Experience helped. It did not protect. A human in the loop is not an independent check unless the workflow is designed to make it one, and Annex 22 does not tell you how.
Qualification is a compound key
The commonest design error is treating operator qualification as a property of the person. It is a property of a tuple: this person, on this system, on this model version, for this use case.
That is not a clause you can quote. It falls out of reading clause 3.3 alongside clause 10.1's change-control scope, which covers the model, the system and the whole process the model assists, and alongside the binding requirement in EU GMP Chapter 2 clause 2.11 that continuing training be given and its "practical effectiveness should be periodically assessed", with records kept.
| Component that changes | What lapses | Typical trigger nobody logs |
|---|---|---|
| The person | Everything | Reviewer leaves the team, covers a shift, returns from twelve months away |
| The model version | Performance evidence for the pairing | Vendor pushes a routine update |
| The use case | Intended-use description and operator responsibility | Same model pointed at a second product family |
| The interface | Review behaviour, though not model performance | UX change that surfaces a confidence score differently |
The third row quietly breaks compliance programmes. A model qualified for deviation triage at one site is extended to a second with different equipment, and the qualification record follows the model rather than the pairing. Annex 22 clause 3.2 asks for the input sample space to be split into subgroups by characteristics including "geographical site or equipment" for exactly this reason.
The joint EMA–FDA guiding principles of good AI practice in drug development of January 2026 make the same point from the other direction: Principle 8 says risk-based performance assessments "evaluate the complete system including human-AI interactions". Note also what that two-page document does not do. Searching it on 30 August 2026, "oversight" appears once, in Principle 2, and "human oversight" and "human-in-the-loop" do not appear at all. It states shared direction, not requirements.
The 2 numbers that separate a control from a slogan
If clause 3.3 requires the operator's consistent performance to be monitored, you need a metric. Two are cheap, are ordinary telemetry, and survive contact with an inspector.
Override rate by output type. The proportion of outputs the reviewer rejects or modifies, segmented by kind of output. A rate near zero says the reviewer is rubber-stamping. A rate near 100% says the model is not fit for the intended use and the human is doing the work unaided. Neither reading is available if the system records only the final approved value.
The clinical decision support literature shows why the segmentation matters. Studying inpatient medication alerts in JAMIA, Nanji and colleagues found 73.3% overridden with about 60% of overrides appropriate — but appropriateness ranged from 2.2% for renal-based substitutions to 98% for duplicate drugs. Their outpatient study found 52.6% overridden, 53% appropriate, ranging from 12% to 92% by alert type. The aggregate number tells you almost nothing. The segmented one tells you which output types your reviewers can actually judge.
Median time to decision. The interval between an output being presented and a decision being recorded, tracked as a distribution rather than a mean. The tail is the finding: a cluster of sub-two-second approvals on a task that takes a competent person forty seconds is documentary evidence that the control was not operating, and it sits in your audit trail whether or not you look at it first.
Neither metric is named in Annex 22. Both are the most defensible way to satisfy the clause 3.3 monitoring duty, and both are trivial to instrument at build time and expensive to retrofit. In work such as LLM validation under GxP, where the human signature is often the whole compliance argument, the same two numbers are what distinguish a reviewed workflow from an unreviewed one.
What an inspector can already cite, without waiting for Annex 22
None of this needs a new rule to bite. On 2 April 2026 the FDA issued a warning letter to Purolea Cosmetics Lab of Livonia, Michigan, after an inspection run from 28 to 30 October 2025. The firm had used AI agents to draft drug product specifications, procedures and master production records, then used those outputs without further review. FDA cited 21 CFR 211.22(c), which requires the quality control unit to approve or reject "all procedures or specifications impacting on the identity, strength, quality, and purity of the drug product", plus 21 CFR 211.100 for absent process validation. The firm's explanation for the missing validation was that the AI agent had not told them it was required.
That rule comes from the 1978 CGMP regulations for finished pharmaceuticals and says nothing about AI. It did not need to. Our account of what the Purolea letter cited and why the age of the rule matters sets out the enforcement logic: missing documented human approval is already an inspectable deficiency in the United States.
What this means in practice
Start by finding the sentence. Search your AI risk assessments and validation plans for anywhere that reduced testing, a smaller test set or a lower criticality classification is justified by human review. Every hit is a place you have taken on the clause 3.3 obligation, knowingly or not.
For each one, do four things. Write the operator's responsibility into the intended-use document rather than an SOP appendix, because that is where clause 3.3 puts it and that is what a process SME approves before acceptance testing starts. Instrument override rate and time-to-decision before go-live, segmented by output type, with pre-agreed action thresholds; decide now what an override rate of 2% triggers, because you will see one. Define the qualification tuple explicitly and wire it to change control, so a vendor model update raises the question of whether reviewer qualification still holds. And establish the human-only baseline before you deploy, because clause 4.3 will ask for it and reconstructing it afterwards is close to impossible.
The third of those is the expensive one, because most LMS configurations cannot represent a compound key. Budget a fraction of an FTE per system for monitoring and periodic reassessment, plus whatever your quality unit charges for a new record type. Against the model testing it buys, that is usually still the better trade. It is a bad trade only when nobody priced it, the reviewers were never qualified as a monitored process, and the first person to notice is an inspector reading an audit trail full of two-second approvals.
Questions people ask about this
- Does human review reduce how much AI validation I have to do?
- Under draft EU GMP Annex 22, published for consultation on 7 July 2025 and not yet binding, it can — but the trade is explicit. Clause 3.3 says that where testing effort has been diminished because a human operator reviews the output, the training and consistent performance of that operator must be monitored like any other manual process. You move cost rather than remove it.
- Is Annex 22 law in August 2026?
- No. On 30 August 2026 the binding EU GMP text for computerised systems is the 2011 version of Annex 11. Draft Annex 22 and the draft Annex 11 revision were published for public consultation on 7 July 2025, the consultation closed on 7 October 2025, and no final text has been adopted. The EMA inspectors working group has been targeting Q4 2026 for delivering final text.
- What is automation bias and does pharma regulation name it?
- Automation bias is the tendency to accept a system output without independent checking. EU AI Act Article 14(4)(b) names it directly, requiring that oversight persons remain aware of the tendency to over-rely on high-risk system output. The 7 July 2025 draft of Annex 22 does not use the term anywhere in its six pages, so pharma teams have to import the concept themselves.
- What should I measure to prove human oversight is working?
- At minimum, the reviewer override rate broken down by output type, and the median time between an output being presented and a decision being recorded. Both are ordinary system telemetry. Without them you cannot show an inspector that review is happening, that it is discriminating between good and bad outputs, or that performance has stayed consistent since qualification.
- Can generative AI be used in GMP if a human checks the output?
- Draft Annex 22 states that generative AI and large language models should not be used in critical GMP applications at all. For non-critical applications it says personnel with adequate qualification and training should always be responsible for ensuring outputs are suitable for the intended use. The human step is what keeps the use non-critical, so it has to be real.