Validating AI Under GxP in 2026: What Binds You, What Is Draft, What Inspectors Cite
There is no AI validation regime. There is ordinary computerised-system law, a six-page draft supplement that is not yet adopted, and one American enforcement action that applied a rule written in 1978. Knowing which is which is the whole job.
On 30 August 2026 the operative computerised-systems text in EU GMP is still Annex 11 as it came into operation on 30 June 2011. It runs to five pages, carries seventeen numbered clauses, and contains no occurrence of the words "artificial intelligence" or "machine learning". Every AI system running in a European GMP process today is validated against it.
That is the whole answer to the question most quality directors are actually asking, and it survives one level down. The revised Annex 11 published for consultation on 7 July 2025 grew to nineteen pages and seventeen chapters. Searched in full text on 30 August 2026, the version downloaded from the Commission's consultation page mentions artificial intelligence exactly zero times, mentions machine learning zero times, and never once refers the reader to Annex 22. Its only cross-reference to another annex is to Annex 15 on qualification and validation. The AI layer of European GMP is a separate six-page supplement, bolted on top, cross-referencing Annex 11 twice while Annex 11 does not acknowledge it exists.
None of it is law yet. The one place where a regulator has actually acted against an AI-generated GxP record is the United States, and the citation was 21 CFR 211.22(c), a clause promulgated in the 1978 CGMP final rule at 43 FR 45077 and unchanged in substance since. The rule that caught AI was written before the personal computer.
- On 30 August 2026 the binding EU computerised-systems text is Annex 11 (2011). Draft Annex 22 and the draft Annex 11 revision are consultation documents from 7 July 2025 and have not been adopted.
- The 19-page draft Annex 11, searched in full on 30 August 2026, contains zero mentions of artificial intelligence and no cross-reference to Annex 22. The AI layer is a separate six-page annex.
- FDA has already enforced against AI-generated GMP documents using 21 CFR 211.22(c), in a warning letter to Purolea Cosmetics Lab dated 2 April 2026. No new rule was needed.
- On the GCP side the pattern repeats: the EMA computerised-systems guideline effective 9 September 2023 explicitly defers AI-specific requirements to a future annex.
- The EU AI Act touches you through Article 50 transparency (live 2 August 2026) and the AI literacy duty, not through the high-risk regime, which moved to 2 December 2027 under Regulation (EU) 2026/1744.
What is actually binding on 30 August 2026
The word "binding" needs care in this domain, because EudraLex Volume 4 is formally guidance. The cover page of the 2011 Annex 11 states its own legal basis: Article 47 of Directive 2001/83/EC, and it "provides guidance for the interpretation of the principles and guidelines of good manufacturing practice (GMP) for medicinal products as laid down in Directive 2003/94/EC". That directive was itself repealed by Commission Directive (EU) 2017/1572, with references now read across to 2017/1572 and to Delegated Regulation (EU) 2017/1569 for investigational products. The veterinary limb it cites, Directive 2001/82/EC, was replaced by Regulation (EU) 2019/6. The text that governs every GMP computerised system in Europe still points at two repealed directives on its front page.
Practically, none of that gets you out of anything. A manufacturing authorisation holder must comply with GMP, and inspectors read Annex 11 as the statement of what compliance means. Treat it as binding. Treat the July 2025 drafts as forecasting.
| Instrument | Status on 30 Aug 2026 | Key date | What it says about AI |
|---|---|---|---|
| EU GMP Annex 11 (2011) | Operative guidance, inspected against | Came into operation 30 Jun 2011 | Nothing. No AI or ML mention in 5 pages |
| Draft revised Annex 11 | Consultation draft, not adopted | Published 7 Jul 2025, consultation closed 7 Oct 2025 | Nothing. Zero AI mentions in 19 pages |
| Draft Annex 22 (AI) | Consultation draft, not adopted | Published 7 Jul 2025, closed 7 Oct 2025 | The entire AI layer, 6 pages |
| 21 CFR 210/211 | Binding US regulation | 211.22 from 43 FR 45077, 29 Sep 1978 | Nothing AI-specific. Already used to cite AI misuse |
| 21 CFR Part 11 | Binding, with 2003 enforcement discretion | Guidance issued 2003 | Nothing AI-specific |
| EMA computerised systems in clinical trials | Guideline in effect | Effective 9 Sep 2023 | Defines AI, then explicitly defers requirements |
| EU AI Act, Reg. 2024/1689 as amended | Binding regulation, phased | Art. 50 live 2 Aug 2026 | Horizontal, not GxP-specific |
The consequence for a validation plan is unglamorous. In August 2026 you cannot write "in accordance with Annex 22" in a validation summary report, because there is nothing to be in accordance with. You write your rationale against Annex 11 and your quality system, and you note where you have anticipated the draft. Auditors accept that. What they do not accept is a plan that cites a draft clause number as though it were a requirement, because that tells them you have not read the source.
The draft Annex 11 does not mention artificial intelligence
This deserves stating precisely, because it contradicts a great deal of published commentary. Several consultancy summaries describe the revised Annex 11 as extending to AI and machine-learning systems. The document does not.
The version searched is the consultation draft published on 7 July 2025 and hosted on the Commission's stakeholder consultation page as mp_vol4_chap4_annex11_consultation_guideline_en.pdf, downloaded and searched in full text on 30 August 2026. It is nineteen pages. Its seventeen chapters are Scope, Principles, Pharmaceutical Quality System, Risk Management, Personnel and Training, System Requirements, Supplier and Service Management, Alarms, Qualification and Validation, Handling of Data, Identity and Access Management, Audit Trails, Electronic Signatures, Periodic Review, Security, Backup and Archiving. Searches for "artificial intelligence", "machine learning" and the standalone token "AI" return nothing. A search for cross-references to other annexes returns one hit, to Annex 15.
The relationship runs one way. Annex 22's scope paragraph says it "provides additional guidance to Annex 11 for computerised systems in which AI models are embedded", and its clause 4.3 refers the reader to "Annex 11 2.7" for the principle that a model must perform at least as well as the process it replaces. In the published draft Annex 11, clause 2.7 is Security. The no-decrease principle is at 2.8. The AI annex's only substantive cross-reference into its parent document points at the wrong clause, which is the sort of thing that gets fixed between consultation and adoption and is worth watching as a tell for how tightly the two texts were drafted together.
The design decision underneath is deliberate and defensible. Annex 11's chapters do the heavy lifting for any system, AI or not: risk management, supplier management, qualification, data handling, access control, audit trails, periodic review. Clause 2.5 says system requirements "should serve as the very basis for system qualification and validation". Clause 2.8 says a replacement system must not decrease product quality, patient safety or data integrity, or increase overall process risk. Those two sentences already dispose of most bad AI proposals. If you cannot write down what the model is for and cannot show the process is no worse with it, you have failed Annex 11 before anyone opens the AI annex. A clause-by-clause reading of what the new AI annex adds on top of ordinary computerised-system requirements is the second document to read, not the first, and the seventeen chapters of the Annex 11 revision that will consume budget are where the money actually goes.
Why the AI layer is only six pages
Annex 22 is short because it does one job: it adds model-specific expectations to systems that Annex 11 already governs. Ten sections, six pages, a glossary. Scope, Principles, Intended Use, Acceptance Criteria, Test Data, Test Data Independency, Test Execution, Explainability, Confidence, Operation.
Its scope section is where the industry argument lives. The draft applies to machine-learning models that "obtained their functionality through training with data, rather than being explicitly programmed". It then narrows twice. It applies to static models only, and says dynamic models that continuously learn during use "should not be used in critical GMP applications". It applies to deterministic models only, and says models with probabilistic output "should not be used in critical GMP applications". Then it draws the conclusion: "the document does not apply to Generative AI and Large Language Models (LLM), and such models should not be used in critical GMP applications."
Read the next sentence, which most summaries omit. For non-critical applications, the draft says personnel with adequate qualification and training "should always be responsible for ensuring that the outputs from such models are suitable for the intended use, i.e. a human-in-the-loop (HITL)". That is not a ban on generative AI in manufacturing. It is a ban on generative AI as the decision-maker where the decision touches patient safety, product quality or data integrity, and a competence-and-review requirement everywhere else.
The technical requirements are stricter than most validation teams expect, and they are stricter about people than about models. Clause 6.5 requires procedural or technical controls preventing staff who have seen the test data from working on training and validation of the same model, and where an organisation cannot maintain that separation it requires a four-eyes arrangement with a colleague who has not had access. Clause 6.2 requires that if test data is split from a pool before training, the developers must never have had access to it, that the test data sits behind access control and audit trail, and that no copies exist outside that repository. Clause 5.6 says generating test data or labels with generative AI "is not recommended" and requires full justification. Clause 8.1 requires feature attribution such as SHAP or LIME, or visual tools such as heat maps, captured during testing of critical models. Clause 9.1 requires the system to log a confidence score for each prediction, and 9.2 requires a threshold below which the model returns "undecided" rather than a possibly unreliable answer.
There is one more detail worth noticing. The Annex 22 glossary defines "AI system" in the exact words of Article 3(1) of the EU AI Act: a machine-based system designed to operate with varying levels of autonomy, that may exhibit adaptiveness after deployment, and that infers from input how to generate outputs such as predictions, content, recommendations or decisions. GMP has adopted the horizontal regulation's definition wholesale. Whatever counts as an AI system for the AI Act will count as one for Annex 22, which matters when you are arguing that a rules engine with a regression fit is not in scope.
The draft is also under active reconsideration. EMA held a multistakeholder expert workshop on 30 June and 1 July 2026 specifically to gather evidence on control measures and guardrails, after consultation feedback showed support for enabling generative models under a risk-based approach. The open questions listed for that workshop include whether adaptive and probabilistic models can be accommodated at all, whether guardrails reliably prevent fabrication, and whether certain critical GMP functions should stay off-limits. Anyone who tells you the final text will match the July 2025 draft is guessing.
What FDA has already cited
The American position needs no forecasting, because it has been demonstrated.
On 2 April 2026 FDA issued a warning letter to Purolea Cosmetics Lab (letter number 722591-04022026). The firm had used AI agents to create drug product specifications, procedures, and master production and control records, and used the resulting documents without adequate quality-unit review. FDA cited 21 CFR 211.22(c), the clause requiring that the responsibilities and procedures applicable to the quality control unit be in writing and followed, and that the unit approve or reject procedures and specifications affecting identity, strength, quality and purity. The letter also cited 21 CFR 211.100 for a process-validation failure, in circumstances where the firm reportedly told investigators it had not known validation was required because the AI agent had not said so. The products were held adulterated under section 501(a)(2)(B) of the Federal Food, Drug, and Cosmetic Act.
Two things follow. First, no new authority was needed. Section 211.22 traces to the 1978 CGMP final rule at 43 FR 45077, published 29 September 1978. The agency did not reach for the January 2025 AI draft guidance or for any AI-specific instrument. It reached for the oldest applicable rule about who signs. The enforcement action that applied a 1978 quality-unit rule to AI-generated documents is the clearest available statement of how this will be inspected: accountability does not transfer to the tool.
Second, the firm was small and the facts were extreme. Do not over-read a single letter as a trend, and be careful with the superlatives attached to it in trade coverage. What it reliably establishes is the citation pathway an investigator will use, not the frequency with which they will use it.
For US drug sponsors there is a second trap in the neighbourhood. FDA finalised its Computer Software Assurance guidance on 24 September 2025, and its philosophy of risk-based, unscripted testing is genuinely convergent with GAMP 5 Second Edition. Its scope is medical-device production and quality-system software under Part 820, now the QMSR. Citing it as the governing framework for a drug sponsor's GCP or GMP system is a specific, checkable error, and the reasons computer software assurance does not cover a drug sponsor's systems are worth knowing before you put it in a validation plan.
Does the EU AI Act change your validation plan?
Less than the volume of commentary suggests, and not in the way most people assume.
The high-risk regime is the part everyone talks about and the part least likely to catch a pharmaceutical process. The Commission's Digital Omnibus on AI became Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026 and in force from 27 July 2026, six days before the original deadline. It moved obligations for standalone Annex III high-risk systems from 2 August 2026 to 2 December 2027, and for AI embedded in products already covered by EU product-safety law, Annex I, to 2 August 2028. That second date is the one that matters for AI in medical devices under the MDR and IVDR. It does not generally catch a deviation-triage model in a fill-finish plant, because Annex III targets employment, credit, biometrics, essential services and similar categories.
What is live is more mundane and more likely to apply to you. Prohibited practices and the Article 4 AI literacy duty have applied since 2 February 2025. GPAI model obligations have applied since 2 August 2025. Article 50 transparency became generally applicable on 2 August 2026 and was deliberately left out of the deferral: where a system interacts with a person or generates synthetic content, that has to be disclosed and machine-readably marked, with providers of generative systems already on the market before 2 August 2026 given until 2 December 2026 to meet the marking obligation.
For a GxP AI deployment this means the AI Act shows up in your documentation as a transparency and literacy obligation, sitting beside your validation package rather than inside it. The training records your quality system already keeps will carry most of the Article 4 burden. The disclosure requirement is a user-interface decision, and it is cheap if you make it at design time and expensive if you retrofit it across twelve validated screens.
The clinical side deferred AI the same way
The GMP pattern is not an accident of one drafting group. It repeats in GCP.
The EMA guideline on computerised systems and electronic data in clinical trials, EMA/INS/GCP/112288/2023, was adopted by the GCP Inspectors Working Group on 7 March 2023 and came into effect six months after publication, on 9 September 2023. It is fifty-two pages, it introduced ALCOA++ with traceability as an explicit attribute, and it does discuss AI. It defines artificial intelligence in its glossary and lists AI among the technologies driving the guideline.
Then, in its scope section, it says this about AI used for participant recruitment, eligibility determination, coding of events and concomitant medication, data clarification, query processes and event adjudication: "Requirements to AI beyond the generally applicable expectations to all systems will not be covered in this guideline initially. This may be covered in a future Annex."
Read that alongside Annex 11 and the structure of European regulatory thinking is visible. AI is validated as a computerised system under the ordinary rules, and the AI-specific layer is deferred to a separate annex that does not yet exist in GCP and is not yet adopted in GMP. Which is a coherent position, and a considerably more demanding one than it sounds, because the ordinary rules include user requirements, risk assessment, supplier assessment, traceability, audit trail and periodic review for a system whose behaviour you cannot fully specify in advance. Getting to a defensible answer on validating a model that gives different answers to the same input is where most of the actual engineering effort goes.
The guidance that is not law but decides your inspection
Below the binding layer sits a set of documents with no legal force that will nonetheless shape what an inspector expects to see. Getting their status right in a validation plan is a two-minute job that buys credibility for the rest of the document.
| Document | Publisher | Date | Status |
|---|---|---|---|
| GAMP 5 Second Edition | ISPE | July 2022 | Industry guidance, not regulation |
| GAMP Guide: Artificial Intelligence, 290 pages | ISPE | July 2025 | Industry guidance; companion to GAMP 5 2nd Ed |
| Considerations for the Use of AI to Support Regulatory Decision-Making | FDA | Draft, 6 Jan 2025 | Draft guidance, non-binding, not finalised as of Aug 2026 |
| Reflection paper on AI in the medicinal product lifecycle | EMA | 9 Sep 2024 | Reflection paper, no legal force |
| Guiding principles of good AI practice in drug development | EMA and FDA jointly | January 2026 | Ten principles, aspirational, no legal force |
| AI in pharmacovigilance, CIOMS Working Group XIV | CIOMS | Final report, December 2025 | Consensus report; the consultation draft was 1 May 2025 |
Two of these are worth using rather than merely citing. FDA's January 2025 draft sets out a seven-step credibility assessment for a model in a defined context of use, with model risk framed as model influence multiplied by decision consequence. It remains a draft on 30 August 2026 and is marked "not for implementation", but its structure is the cleanest available way to write down why you tested what you tested. The joint EMA and FDA guiding principles of January 2026 give you ten short statements that both regulators have signed, including risk-based validation proportionate to context of use, clear context of use, life-cycle management with scheduled monitoring for data drift, and risk-based performance assessment that evaluates "the complete system including human-AI interactions". That last phrase is the one to quote at anyone proposing to test the model in isolation and call it validated.
The CIOMS Working Group XIV report on AI in pharmacovigilance is a final consensus report published in December 2025, not a draft, and it is the document a PV head will be measured against even though it binds nobody. Cite the final report, not the 1 May 2025 consultation draft, and say which you mean.
What this means in practice
Start with the honest sentence about status, because it is the cheapest credibility you will ever buy. In your validation plan, write that the applicable requirements are Annex 11 (2011) and your quality system, that Annex 22 and the revised Annex 11 are consultation drafts published 7 July 2025 with consultation closed 7 October 2025 and not adopted as of the plan date, and that the design has been assessed against the draft expectations as forward-looking risk mitigation. Three sentences. They pre-empt the entire "why did you not follow Annex 22" line of questioning and they signal that you read the sources.
Then classify by criticality, not by technology. The Annex 22 test is functional: does the output have direct impact on patient safety, product quality or data integrity? A model that ranks a deviation queue for a qualified person who assigns severity is not making the decision. A model whose output is the disposition is. Most deployments can be designed onto the safe side of that line at negligible cost if the decision is made before the architecture is fixed, and at very high cost afterwards. This is a design review item, not a compliance review item.
Build the human step so it produces evidence. The Purolea citation was not about AI quality. It was about a quality unit that did not review. If a qualified person reviews model output, the record must show who reviewed, what they saw, what they changed, and when. Annex 22 clause 10.5 anticipates exactly this and warns that where testing effort has been reduced because a human is in the loop, the review may need to cover every output. That trade is real: you either test the model hard or you review its output hard, and you should decide which deliberately rather than discovering it during an inspection.
Budget the test data, not the model. The expensive clauses in Annex 22 are 5.x and 6.x. Test data that is stratified across subgroups, sufficient in size for statistical confidence, labelled to a very high degree of correctness through independent verification, held behind access control with an audit trail, and never touched by the people who trained the model. In a mid-size manufacturing deployment that data work, and the staff-separation controls around it, will typically outweigh model development effort. The GAMP community's own published estimate is that validation should account for roughly ten per cent of overall project budget when planned early. Treat that as a floor for AI systems, and treat the widely circulated vendor figures for validation-effort savings as marketing rather than benchmarks, because most of them trace back to a single unsourced slide.
Name the signatories before you build. For a GMP AI system that is the process subject matter expert who owns the intended-use description and the acceptance criteria (Annex 22 puts this on the SME explicitly, in clauses 3.1 and 4.2, and requires approval before acceptance testing starts), the system owner, and QA. If those three people cannot be named in the first workshop, the project is not ready to start.
Finally, keep a watch item on the EMA calendar rather than a plan built on a date. The consultation closed on 7 October 2025 and drew roughly 1,300 comments. EMA ran an expert workshop on 30 June and 1 July 2026 on whether guardrails can make probabilistic models acceptable. Final texts have been signalled for the course of 2026 with operative dates to be confirmed, and that signalling has already slipped once. Design so that a relaxation of the generative-AI position is upside rather than rework, and so that adoption of the draft as written costs you nothing.
Questions people ask about this
- Is Annex 22 law?
- No. On 30 August 2026 Annex 22 is a consultation draft. The European Commission published it for stakeholder comment on 7 July 2025 and the consultation closed on 7 October 2025. No final text has been adopted. The binding computerised-systems guidance in EU GMP remains Annex 11 as it came into operation on 30 June 2011.
- What regulation actually applies to an AI system in a GxP process today?
- The same rules that apply to any computerised system in that process. In EU GMP that is Annex 11 (2011) read with Chapter 4 and the GMP directives. In US drug manufacturing it is 21 CFR Parts 210 and 211, with Part 11 for electronic records. In EU clinical trials it is the EMA computerised-systems guideline effective 9 September 2023. None of these contains AI-specific requirements.
- Has FDA ever cited a company for using AI in a GxP process?
- Yes. A warning letter to Purolea Cosmetics Lab dated 2 April 2026 cited 21 CFR 211.22(c) because the quality unit did not review documents generated by AI agents for accuracy and CGMP compliance, and 21 CFR 211.100 for a process-validation failure. FDA applied existing CGMP rules rather than any AI-specific requirement.
- Does the EU AI Act apply to pharmaceutical AI systems?
- Partly, and mostly not through the high-risk regime. Prohibited practices and the Article 4 AI literacy duty have applied since 2 February 2025, GPAI obligations since 2 August 2025, and Article 50 transparency since 2 August 2026. High-risk obligations for Annex III systems were moved to 2 December 2027 by Regulation (EU) 2026/1744, which entered into force on 27 July 2026.
- Can I validate a large language model for a GxP process?
- You can validate the system it sits inside, for a narrowly defined context of use, with human review as a designed control. What you cannot do today is put a generative model in the critical decision path in EU GMP: the draft Annex 22 states that generative AI and large language models should not be used in critical GMP applications, and inspectors already read the draft even though it is not adopted.