GxPGxP, validation & assurance

EU GMP Annex 22, Clause by Clause

The draft annex is six pages long and most of it is unremarkable. Three clauses are not, and one of them quietly turns every GMP AI project into two projects.

Draft Annex 22 is six pages and roughly 2,300 words. Ten numbered sections and a glossary, most of it reading like a competent validation engineer writing down what they already do. Then you reach clause 4.3, which is two sentences long and reorders the budget of every AI project in a GMP environment.

The clause says the acceptance criteria of a model "should be at least as high as the performance of the process it replaces", then adds the sentence that does the damage: "This implies, that the performance should be known for the process which is to be replaced by a model." Nobody argues with the first sentence. The second obliges you to produce a number that, for most of the processes people want to automate, has never been measured. A deviation-triage model has to beat a triage process whose accuracy nobody has quantified. A model reading batch records has to beat reviewers whose miss rate has never been sampled. Before you can validate the model, you have to characterise the human.

Everything below is read against the consultation text the European Commission published on 7 July 2025. It is a draft. The consultation closed on 7 October 2025 and, as of 30 August 2026, no final version has been adopted — the binding EU requirement for computerised systems in GMP is still the 2011 Annex 11. Designing to the draft is still the right call, and the clause-by-clause reading below is why the expensive clauses stay expensive whether or not they become law.

In short
  • Annex 22 remains a consultation draft on 30 August 2026. Ten sections, six pages, no legal force. The 2011 Annex 11 binds.
  • Clause 4.3 requires acceptance criteria at least as high as the replaced process, and states the replaced process's performance must therefore be known. That is a separate measurement study, before the model project starts.
  • Clause 5.6 discourages generating test data or labels with generative AI. Clause 6.2 requires that staff who trained the model never had access to the test data, with a four-eyes fallback in 6.5.
  • Clause 6.2 is the one that invalidates existing pilots: it is retrospectively unprovable for almost every model built on a random train/test split.
  • Clause 8.1 names SHAP, LIME and heat maps in the text. Clause 9.2 asks for an "undecided" output. Both are design constraints, not documentation.

What is binding on 30 August 2026, and what is not

The distinction matters more than the content, because a consultant who cites Annex 22 as a requirement has told the room they have not read the front page.

InstrumentStatus on 30 August 2026Scope
EU GMP Annex 11 (January 2011)BindingComputerised systems: validation, data, audit trail, security, e-signatures
Draft revised Annex 11 (7 July 2025)Draft; consultation closed 7 Oct 202519 pages, 17 chapters. No AI content at all
Draft Annex 22 (7 July 2025)Draft; consultation closed 7 Oct 2025AI/ML models in critical GMP applications
Draft revised Chapter 4 (7 July 2025)Draft; consultation closed 7 Oct 2025Documentation, hybrid records
Regulation (EU) 2026/1744 (AI Omnibus)In force since 27 July 2026Amends the AI Act; defers high-risk dates

All three GMP drafts appeared together on the Commission's stakeholder consultation page, which closed at 23:59 CEST on 7 October 2025 and publishes no comment tally. Anyone quoting a precise number of consultation responses is quoting a secondary source.

Two things are worth knowing about the timetable. The EMA GMP/GDP Inspectors Working Group work plan for 2024–2026 targets Q1 2026 for delivering final texts of Chapter 4 and Annex 11 to the European Commission. That quarter passed with nothing published. The same work plan has no Annex 22 line item, because the annex did not exist when it was written — so the widely repeated "Q4 2026" date for Annex 22 does not come from it.

The position in the draft is also moving. EMA held a multistakeholder workshop on 30 June and 1 July 2026, an open expert session followed by a closed drafting-group session, to gather evidence on control and mitigation measures that might admit dynamic, adaptive and probabilistic models, generative AI included, under a risk-based approach. The scope exclusion in the July 2025 draft is under review by the people who wrote it. The evidence requirements in sections 4 to 10 are not.

Annex 22 is also not the EU AI Act. The AI Omnibus, Regulation (EU) 2026/1744, entered into force on 27 July 2026, deferring standalone Annex III high-risk obligations to 2 December 2027 and product-embedded Annex I obligations to 2 August 2028; prohibited practices and the Article 4 AI literacy duty have applied since 2 February 2025 and were not deferred. Most AI inside a GMP process is not high-risk under the AI Act at all, and is in scope of GMP regardless. For the wider map, see what binds you, what is draft and what inspectors cite in GxP AI validation.

Sections 1 to 3: scope, and the bargain in clause 3.3

Section 1 draws the boundary functionally: models used "in critical applications with direct impact on patient safety, product quality or data integrity", which "obtained their functionality through training with data, rather than being explicitly programmed". Only static models with deterministic output are covered. Dynamic models that learn during use, and probabilistic models that may return different outputs for identical inputs, fall outside and "should not be used in critical GMP applications". Generative AI and LLMs follow from those two exclusions rather than being banned separately.

Section 2 is three clauses on personnel, documentation and quality risk management. Clause 2.2 matters most to anyone buying rather than building: documentation "should be available and reviewed by the regulated user" whether the model was trained in-house or supplied by a vendor. A supplier who will not hand over training and test documentation has made your model unqualifiable, and no contract clause fixes that after signature.

Section 3 writes the intended use. Clause 3.1 wants "a comprehensive characterisation of the data the model is intended to use as input and all common and rare variations" — the input sample space — with limitations and biased inputs identified, approved by a process subject matter expert before acceptance testing starts. Clause 3.2 divides that space into subgroups: by decision output, by site or equipment, by material, by defect type and severity.

Clause 3.3 is the human-in-the-loop bargain, and not the escape hatch people read it as. Where a model feeds a human decision "and where the effort to test such model has been diminished", the intended-use description must state the operator's responsibility, and "the training and consistent performance of the operator should be monitored like any other manual process." Clause 10.5 completes the trade: records of that review must be kept, and depending on criticality it "may imply a consistent review and/or test of every output from the model". You spend less on testing the model and spend it qualifying the person instead, which is the case for treating human oversight as a control you qualify rather than a sentence in a risk assessment.

Section 4: the clause that turns one project into two

Clause 4.1 asks for case-dependent test metrics and offers examples for a classifier: confusion matrix, sensitivity, specificity, accuracy, precision, F1. Clause 4.2 requires acceptance criteria against those metrics, approved by a process SME before testing, and permits different criteria for different subgroups. Both are ordinary.

Clause 4.3 is titled "No decrease" and reads in full: "The acceptance criteria of a model, should be at least as high as the performance of the process it replaces. This implies, that the performance should be known for the process which is to be replaced by a model (see Annex 11 2.7)."

Read that cross-reference literally and it does not resolve. In the 7 July 2025 consultation draft of Annex 11 — the 19-page PDF from the Commission's consultation page, searched on 30 August 2026 — clause 2.7 is "Security", about keeping up to date with security threats. The clause carrying the no-decrease principle is 2.8, "No risk increase": where a computerised system replaces another system or a manual operation, "there should be no resultant decrease in product quality, patient safety or data integrity. There should be no increase in the overall risk of the process." The pointer is off by one, almost certainly an artefact of the two drafts being renumbered before publication, and presumably will be corrected. The principle is unambiguous either way, and it long predates AI.

What AI adds is the word "known". Annex 11's 2.8 prohibits getting worse. Annex 22's 4.3 says you must be able to prove you have not, which requires a number for the thing being replaced. For most candidate processes that number does not exist.

Process being replacedBaseline method established?What measuring it takes
Manual visual inspection of injectablesYes — probability of detection on seeded defect setsReference inspector panels; defect kits already exist
Deviation severity triageNoBlinded re-adjudication of a sampled deviation set
Batch record review for errorsNoSeeded-error study, or blinded double review of a sample
Audit trail anomaly reviewNoNeeds a labelled set of true anomalies, which is the hard part
Environmental monitoring plate readingPartly — method comparison practice existsSplit-sample study against reference readers

Parenteral visual inspection is the outlier, and instructive because it is the exception. USP General Chapter <1790> treats visual inspection as a probabilistic process — "the specific detection probability observed for a given product for visible particles will vary" with dosage form, particle characteristics and container design — and the field qualifies alternatives to manual inspection relative to measured manual performance rather than against an absolute standard, using the probability-of-detection approach that traces to Knapp and Kushner in 1980. That discipline has had an answer to clause 4.3 for forty-five years. Nobody else has.

Two consequences surprise steering committees.

The baseline study lives in the quality system. It has a protocol, a sample size, a blinding scheme and a report, and its output is a documented statement about how well your people currently perform a GMP-relevant task. If that number is poor, you have not only built a business case for the model. You have generated a finding about an existing process, in a controlled document, that an inspector can ask to see. The correct response is a CAPA, and the project plan should assume one.

Clause 4.2 permits subgroup-specific criteria, so it implies subgroup-specific baselines. If the criteria differ for rare defect types, the baseline must resolve at that level, and rare subgroups are where sample size gets expensive. Clause 5.2 then requires the test set and "any of its subgroups" to be large enough to calculate the metrics with adequate statistical confidence. The rare subgroup drives the design of both phases.

Vendor material routinely quotes a percentage cut in false rejects or a lift in detection rate. Each is a relative number presupposing exactly the baseline clause 4.3 demands, measured on someone else's line. Yours has to be produced in your facility, on your process.

Sections 5 and 6: the test data clauses that cannot be satisfied retrospectively

Section 5 has six clauses. Test data must be representative of and expand the full sample space, stratified across all subgroups and rare variations, with the selection rationale documented (5.1), and sufficient in size per subgroup (5.2). Labelling must be "verified following a process that ensures a very high degree of correctness", by independent verification by multiple experts, validated equipment or laboratory tests (5.3). Pre-processing must be pre-specified with a rationale that it represents intended-use conditions (5.4), and any cleaning or exclusion documented and fully justified (5.5).

Clause 5.6 is short: "Generation of test data or labels, e.g. by means of generative AI, is not recommended and any use hereof should be fully justified." That is a discouragement carrying a justification burden, and the burden is heavier than it looks. Read against 5.3, the logic is plain: every test metric is a function of label correctness, so a synthetic label produces a metric computed against a guess. Read against section 6, a second problem appears — a generative model used to produce test data was itself trained on something, and showing its output is independent of your training set is a research exercise, not a validation deliverable.

Section 6 is headed "Test Data Independency" and it is the section that retires pilots.

  • 6.1 requires controls ensuring test data is not used in development, training or validation, either by capturing it only after training and validation are complete, or by splitting it from the pool before training starts.
  • 6.2 is the hard one. Where test data is split from a pool, "it is essential that employees involved in the development and training of the model have never had access to the test data." The data sits under access control and audit trail, and "there should be no copies of test data outside this repository."
  • 6.3 requires a record of which data was used for testing, when, and how many times — which ends the habit of evaluating on the hold-out set repeatedly while tuning, because the count is now evidence.
  • 6.4 covers physical objects: those used for the final test must not previously have been used to train or validate, unless features are independent. If your defect kit was used in development, you need another one.
  • 6.5 requires controls preventing staff who have had access to test data from training and validating the same model. Where that is impossible, such a person "should only have access to training and validation of the same model when working together (in pair) with a colleague who has not had this access (4-eyes principle)".

The bite is that 6.2 and 6.5 are statements about people and history, not about data. Almost every model in a pharmaceutical company's pilot portfolio was built by a small team who had the whole dataset in a notebook and split it with a random seed. No artefact can retrospectively prove nobody looked. That evidence had to exist before the work started: an access-controlled repository, a named roster of who was walled off from it, retained audit trail records. For many pilots the honest answer is that the model must be retrained on a clean split, with a new team arrangement, before it can be tested at all.

Clause 7.4 is the storage consequence nobody budgets for: test documentation retained along with the intended-use description, the characterisation of test data, "the actual test data, and where relevant, physical test objects", plus access control documentation and audit trail records. Retaining physical test objects means a controlled store for vials or components that must not degrade in a way that changes what they demonstrate.

Sections 7 to 9: testing, explainability and confidence

Clause 7.1 states the objective plainly: show the model is fit for intended use and "generalising well", including detection of over- or underfitting. Clause 7.2 lists mandatory test plan contents — intended-use summary, pre-defined metrics and acceptance criteria, reference to the test data, a test script covering every step, and how the metrics are calculated — approved before testing, with a process SME involved. Clause 7.3 makes any deviation from the plan, failure to meet acceptance criteria, or omission to use all test data a documented and investigated event.

Section 8 departs from the technology-neutral drafting EU GMP annexes usually keep to. Clause 8.1 requires systems to record the features contributing to a particular classification or decision during testing, and says "techniques like feature attribution (e.g. SHAP values or LIME) or visual tools like heat maps should be used" where applicable. Clause 8.2 makes a review of those features "part of the process for approval of test results", so a named human must judge whether the model is deciding on relevant features and sign that judgement into the test approval. This is the clause that catches the model achieving excellent metrics by reading a batch identifier, a timestamp or a lighting artefact.

Section 9 is two clauses and both are architecture rather than paperwork. Clause 9.1 requires the system to log the model's confidence score for each prediction. Clause 9.2 requires a threshold so predictions are made "only when suitable", and says that where confidence is very low it "should be considered whether the model should flag the outcome as 'undecided'". A three-state output, with a routing rule for the third state, is a different design from a binary classifier and has to be specified in requirements. Retrofitting it after integrating a service that returns only a class label reopens qualification.

Section 10: operation, and the change that is not to the model

Clause 10.1 puts the model, the system it is embedded in and "the whole process it is automating or assisting" under change control before deployment. Any change to the model, the system or the process, "including any change to physical objects the model is using as input", must be documented and evaluated for retesting, and any decision not to retest fully justified. A vial supplier change, a new label stock, a relocated lamp — all are changes to the input distribution and all enter change control as such. Clause 10.2 adds configuration control with measures to detect unauthorised change.

Clauses 10.3 and 10.4 split monitoring in two. Model performance against its defined metrics is monitored to detect changes in the computerised system, the draft's own example being deterioration of a lighting condition. Separately, input data is monitored to confirm it remains within the model's sample space, with metrics defined for drift. Those answer different questions, and a plan tracking only output accuracy satisfies 10.3 and not 10.4.

Note what the annex ties back to. Annex 22 says it "provides additional guidance to Annex 11 for computerised systems in which AI models are embedded", yet the 7 July 2025 draft of Annex 11, searched on 30 August 2026, contains no occurrence of "artificial intelligence", "AI" or "machine learning". The dependency runs one way. That draft also uses ALCOA+ in clause 2.4 and defines only ALCOA+ in its glossary, not the ALCOA++ formulation secondary coverage often attributes to it. The Annex 11 clauses that cost money for an AI architecture are covered separately.

What this means in practice

Treat any Annex 22-scope project as two projects with a gate between them, and staff them differently.

Phase one is a manual-process characterisation study. Its deliverable is a measured performance figure for the process the model will replace, broken out by the clause 3.2 subgroups, with a protocol, blinding scheme, sample size justification and approved report. It belongs to the process SME, not the data team — the annex names the SME as accountable in 3.1, 4.2 and 7.2. Budget it as a quality study, expect four to twelve weeks, and write the CAPA path into the plan before you see the result.

Freeze the test data and the staff roster before anyone touches training data. The repository, the audit trail, the no-copies rule and the named list of walled-off people all have to exist first, because 6.2 and 6.3 are evidence about the past. Two days of IT and QA work at the start of a project; an unfixable defect at the end of one.

Write the acceptance criteria and the test plan before the model is good. Clauses 4.2 and 7.2 both require approval before acceptance testing begins, and document dates make the sequence visible. Ask the supplier for clause 9.2's "undecided" output in the RFP, not at URS review.

The clause-level test of whether a project is real is simple. Ask what the current process's error rate is. If the room has no number, the project has not started — it has a business case. And when a regulator does look at an AI-supported quality decision, the questions tend to come from the oldest parts of the rulebook, which is the pattern behind an FDA warning letter over AI use and the decades-old predicate rule it cited.

Peer-reviewed interpretation has begun to appear, including in the PDA Journal of Pharmaceutical Science and Technology, and the drafting group's position on generative models may move after the June 2026 workshop. None of that changes sections 4 to 10. The exclusion in section 1 is a policy question regulators can revisit. Knowing how well your process currently works is not.

Questions people ask about this

Is EU GMP Annex 22 legally binding?
No. Annex 22 was published for stakeholder consultation on 7 July 2025 and the consultation closed on 7 October 2025. As of 30 August 2026 no final text has been adopted, so it is a draft with no legal force. The binding EU requirement for computerised systems in GMP remains the 2011 version of Annex 11.
What does Annex 22 clause 4.3 require?
Clause 4.3 of the July 2025 draft states that a model's acceptance criteria should be at least as high as the performance of the process it replaces, and adds that this implies the performance of the replaced process must be known. In practice that means measuring the existing manual or rules-based process before the model can be accepted.
Does Annex 22 allow generative AI to create test data?
Clause 5.6 of the draft says generation of test data or labels, for example by generative AI, is not recommended and any use of it should be fully justified. It is a discouragement carrying a justification burden rather than an outright prohibition, but the burden is heavy because every test metric depends on label correctness.
What is test data independency under Annex 22?
Section 6 of the draft requires that test data is never used in development, training or validation, that it sits under access control and audit trail with no copies elsewhere, and that staff who saw the test data are kept out of training the same model. Where that separation is impossible, clause 6.5 allows a four-eyes arrangement with a colleague who has not seen the test data.
How does Annex 22 relate to Annex 11?
Annex 22 states that it provides additional guidance to Annex 11 for computerised systems in which AI models are embedded, and cross-references it directly. Both were published for consultation on the same day, 7 July 2025, alongside a revised Chapter 4. An AI system in GMP has to satisfy the computerised-systems annex and the AI annex together.