Drift, Abstention and Audit Trails: The Operating Plan a Deployed GxP Model Needs
Validation produces a number on a hold-out set. Operation produces the evidence that the number still holds, and the two are costed as if they were the same activity.
A validated model is a claim about a hold-out set that was true on the day the report was signed. Everything after that day is an argument that the claim still holds, and the instruments that make the argument are cheap at design time and close to impossible to retrofit.
What binds a European manufacturing site today is short. Annex 11 of the EU GMP guide, the 2011 version, gives change management two lines: "Any changes to a computerised system including system configurations should only be made in a controlled manner in accordance with a defined procedure." Clause 11 adds periodic evaluation, clause 9 a risk-based system-generated record of GMP-relevant changes and deletions. Nothing in it anticipates a system whose behaviour degrades while nobody changes it.
The text that does anticipate it is draft Annex 22, published for stakeholder consultation on 7 July 2025 and closed on 7 October 2025. It is not law: as of 30 August 2026 no final text has been adopted and the 2011 Annex 11 remains binding. Its section 10 is still the best available description of what operating a regulated model looks like, and two clauses carry the cost.
- Draft Annex 22 clause 10.4 requires metrics for drift in the input data, separately from clause 10.3's monitoring of model performance. A plan tracking accuracy alone satisfies one and not the other. Both remain consultation text on 30 August 2026.
- Clause 10.1 extends change control to "any change to physical objects the model is using as input". A vial supplier switch is a change to a validated model.
- The word drift appears once in the draft. Retrain and revalidate appear zero times. The retest decision is a change-control decision with a written justification, not an automatic retraining trigger.
- The 0.10 / 0.25 PSI traffic light has no statistical basis: its distribution depends on sample size and bin count, with no stated type I or II error rate.
- OpenTelemetry's gen_ai and mcp conventions give you an ALCOA++-shaped schema for free, but every attribute is marked Development and content capture is off by default.
What binds you on 30 August 2026, and what does not
Binding in the EU for a GMP computerised system: Annex 11 (2011), clauses 1, 9, 10, 11 and 12, plus Chapter 4 of the GMP guide. Clause 4.11 of Chapter 4 is the one people forget when designing telemetry: batch documentation "must be kept for one year after expiry of the batch to which it relates or at least five years after certification of the batch by the Qualified Person, whichever is the longer".
Draft, not binding: the 7 July 2025 revision of Annex 11 with its 17 sections, and Annex 22 with its 10 sections and six pages. Every clause number below prefixed "draft" comes from those.
Live, and routinely confused with this: Regulation (EU) 2026/1744, the Digital Omnibus on AI, in the Official Journal on 24 July 2026 and in force from 27 July 2026, which moved the AI Act's high-risk dates to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I product-embedded ones. Prohibited practices, Article 4 AI literacy, Article 50 transparency and the general-purpose AI obligations already apply. None of it displaces GMP, and none of it tells you when to retest a classifier.
One fact worth holding before an audit. I text-extracted the 7 July 2025 consultation version of draft Annex 11, 19 pages, on 30 August 2026 and searched it: "artificial intelligence", "machine learning" and "drift" occur zero times, and "change control" occurs once, in the glossary, in no numbered clause. The computerised-systems annex routes AI entirely to Annex 22, and Annex 22 is a draft.
Clause 10.4 is the clause that reorders the monitoring plan
The draft splits operational monitoring into two duties that most plans collapse into one.
Clause 10.3, system performance monitoring. "The performance of a model as defined by its metrics should be regularly monitored to detect any changes in the computerised system (e.g. deterioration or change of a lighting condition)." The metrics are those fixed under clause 4.1, which names a confusion matrix, sensitivity, specificity, accuracy, precision and F1.
Clause 10.4, input sample space monitoring. "It should be regularly monitored whether the input data are still within the model sample space and intended use. Metrics should be defined for monitoring any drift in the input data."
These fail at different times. Performance monitoring needs labels, and labels arrive late or never. Input monitoring needs none, which makes it the earlier warning and the cheaper instrument. The sample space is not a monitoring artefact; it is fixed at the start. Clause 3.1 requires the intended-use description to contain "a comprehensive characterisation of the data the model is intended to use as input and all common and rare variations; i.e. the input sample space", approved by a process subject matter expert before acceptance testing begins. Without it, clause 10.4 has no referent.
Then clause 10.1. Change control covers the model, the system it is implemented in and "the whole process it is automating or assisting", and any change "including any change to physical objects the model is using as input" must be documented and evaluated for retesting, with any decision not to retest "fully justified". A new vial supplier, a relabelled carton, a replaced lamp above an inspection station. In a conventional validation these are process changes with no software consequence. Under 10.1 they are changes to a validated model.
The 6 signals, and the metric behind each
Read against the draft, the operating plan resolves to six instruments. Only the first two are named duties; the rest are how you satisfy them without waiting for labels.
| Signal | What it catches first | Practical metric | Draft anchor |
|---|---|---|---|
| Input distribution drift | Supplier, site, equipment or lighting change | PSI, KS statistic or MMD per feature, with declared bins and window size | 10.4 |
| Output distribution shift | Population change with no feature signal | Predicted-class rates and score histogram versus qualification baseline | 10.3 |
| Abstention rate | Inputs leaving the sample space, before any label exists | Share of outputs flagged "undecided", per subgroup, per shift | 9.1, 9.2 |
| Human override rate | Model failure the metrics cannot yet see | Reject or modify rate segmented by output type | 3.3, 10.5 |
| Realised performance | Actual degradation, late but definitively | Clause 4.1 metrics on a sampled, independently labelled audit stream | 10.3 |
| Configuration integrity | Undocumented change to model, thresholds or serving stack | Hash of the deployed artefact set, diffed against the approved baseline | 10.2 |
The sixth gets skipped, and draft Annex 11 clause 14.2(iii) reaches it from the other direction by requiring undocumented changes to be identified "e.g. by means of configuration auditing". If the model, its thresholds, its preprocessing and its serving image are not hashed against an approved manifest, the validated state exists only in a document.
Subgroups matter across all six. Clause 3.2 requires the input sample space to be divided by characteristics such as decision output, geographical site or equipment, and defect type. Aggregate monitoring on a site running four lines reports green while one line degrades, because the other three dilute it.
Why the PSI traffic light will not survive a competent inspector
Population stability index is the default drift metric in most monitoring stacks, and it arrives with a rule of thumb: below 0.10 little change, 0.10 to 0.25 moderate, above 0.25 significant and action required. That rule came out of credit scoring and has no statistical basis.
Bilal Yurdakul and Joshua Naranjo, in the Journal of Risk Model Validation (volume 14, issue 4, 2020), say so directly: "These benchmarks are used without reference to statistical type I or type II error rates." The operationally important part is their result: the distribution of PSI depends on the two sample sizes and the number of bins, not on the distribution of the monitored variable. A fixed 0.25 cut-off therefore carries a different false-alarm rate for a line producing 200 units a shift than for one producing 20,000, and a different rate again once an engineer changes the binning.
Two consequences. The threshold has to be derived from your own baseline and stated as a false-alarm rate you chose, with the derivation retained the way any acceptance criterion is retained; clause 7.4 already requires test documentation and data characterisation to be kept "similarly to other GMP documentation". And the bin definition and the monitoring window are configuration items under clause 10.2. If an engineer can re-bin a monitor without a change record, its sensitivity is uncontrolled and every green result it has produced is unqualified.
Abstention and override rates are drift detectors, not dashboard decoration
Clause 9.1 requires the system to log a confidence score for each prediction where applicable. Clause 9.2 requires a threshold so predictions are made "only when suitable", and says that where confidence is very low it should be considered whether to flag the outcome as "undecided".
Most implementations treat the undecided rate as a cost line. It is better understood as a label-free drift detector. As inputs move away from the training distribution, confidence falls before any ground truth exists to prove accuracy fell with it, so the abstention rate moves earlier than the clause 10.3 metrics and needs no labelling programme to read. A step change on one subgroup, shift or line is a clause 10.4 finding arriving through the clause 9 instrument.
Override rate is the other free signal, and the draft names the duty without naming the metric. Clause 3.3 says that where a human-in-the-loop arrangement has allowed testing effort to be reduced, "the training and consistent performance of the operator should be monitored like any other manual process", and clause 10.5 requires records from that review. You cannot monitor consistent operator performance without a disagreement rate, and a disagreement rate is a drift signal held by the humans.
Three cautions. It is noisy and biased: reviewing 17 studies in JAMIA in 2006, van der Sijs and colleagues found drug safety alerts overridden in 49% to 96% of cases, with overriding frequently justified, so a high rate is a question rather than a verdict. A falling rate is the dangerous direction: Goddard, Roudsari and Wyatt's systematic review of automation bias, also in JAMIA, documents over-reliance on automated advice and the complacency that follows from insufficient monitoring of it, so a rate drifting towards zero is equally consistent with a better model and with a reviewer who has stopped looking. Segmentation by output type and the time-to-decision distribution separate them. And the signal vanishes if the system stores only the approved value: record the disposition without the model's proposal, its confidence and the reviewer's action, and the two cheapest drift detectors you own are unreadable afterwards.
When does a retest actually become mandatory?
In the 7 July 2025 draft of Annex 22, searched on 30 August 2026, "drift" occurs once, in clause 10.4. "Retrain", "revalidate" and "revalidation" occur zero times. There is no automatic retraining trigger in the text, and no clause telling you to withdraw a model from use when a monitor turns red. What there is instead is clause 10.1's evidentiary burden: every change documented, every change evaluated for retesting, every decision not to retest fully justified. The default is a decision with a written rationale, and the artefact an inspector reads is the justification, not the alert.
| Event | Default posture | What "fully justified" has to contain |
|---|---|---|
| Model artefact or version change | Retest | Nothing short of results on the independent test set |
| Threshold, prompt or preprocessing change | Retest the affected metric | Which metric could move, and evidence it did not |
| Physical input object change (supplier, packaging, fixture, lighting) | Evaluate against the sample space | Whether the new object was represented in the clause 3.1 characterisation |
| Drift alert on an input feature | Investigate, then decide | Subgroup analysis plus a performance check on labelled samples from the alert window |
| Clause 4.1 metric below acceptance criterion | Retest, and act on the process | An impact assessment covering product already released |
| Reviewer population change, or clause 3.3 metrics moving | Requalify the pairing | Operator performance evidence for the current model version |
| No change at all, review interval elapsed | Periodic review | Draft Annex 11 clause 14.2's named inputs, including undocumented changes |
The third row catches organisations out, and it is the practical content of "any change to physical objects the model is using as input". Supplier changes route through GMP change control already. What is new is that the change control now has to reach a data scientist, because the question it asks is whether the new object falls inside a sample space characterised months earlier by somebody else.
The audit trail: OpenTelemetry gives you the schema and none of the record
The reflex when a regulated team instruments an AI system is to design a bespoke audit schema. That is largely wasted validation effort, because a vendor-neutral vocabulary for exactly these events exists, is public, and is what the tooling emits anyway.
The OpenTelemetry GenAI semantic conventions model an agent run as a span tree. gen_ai.operation.name covers chat, embeddings, create_agent, invoke_agent, invoke_workflow and execute_tool. Spans carry gen_ai.provider.name, gen_ai.request.model, gen_ai.response.model, token usage, gen_ai.response.finish_reasons and gen_ai.tool.name with call arguments and result. Message content lives in gen_ai.input.messages and gen_ai.output.messages, and a gen_ai.evaluation.result event carries a score value and label, the natural carrier for a monitoring evaluation. The Model Context Protocol attributes mcp.method.name, mcp.session.id, mcp.protocol.version and mcp.resource.uri sit in the same vocabulary, so a tool call against a regulated data store lands in the same trace as the agent that issued it.
Map that onto ALCOA++ and most of it is already there.
| Attribute | What the trace carries | What it does not |
|---|---|---|
| Attributable | Span identity, service, session, model, provider | The accountable human; a system account is not a person |
| Contemporaneous | Start and end timestamps at event time | Nothing, this one is solved |
| Original | Input and output messages, if content capture is enabled | Status as the record of record rather than a copy |
| Accurate | Finish reasons, tool results, evaluation scores | The reason a value was changed |
| Traceable | Trace and span parentage across agent, tool and MCP hops | Linkage to the released batch record |
| Enduring, Available | Nothing by default | Retention; telemetry backends are costed in days |
Four gaps decide whether this is an audit trail or a debugging aid.
Content capture is opt-in, with three modes: not recorded, recorded on span attributes, or stored externally with a span reference. The third is the right one for GxP, the trace pointing at a record in a controlled store rather than becoming the store. Left at the default, the one artefact that is a GMP record, the model's output, is the one you did not keep.
No reason, no signature. Draft Annex 11 clause 12.2 requires the audit trail to capture the user and role, what changed with old and new value, the timestamp, and that "systems should automatically prompt the user for, and register the reason, why the change was made". No gen_ai attribute carries a reason for change or a signature link. That is application work, and it is where the human decision under clause 10.5 attaches.
Clause 12.1 is written about people. It requires audit trail functionality that "automatically logs all manual user interactions". A model's write to a regulated record is not a manual user interaction, which is exactly why the accountable step must be a human action the system logs, and why that action is what the QP sees under draft clause 12.10.
The vocabulary is unstable. As of mid-July 2026 every gen_ai and mcp attribute, span, metric and event in the registry carries the stability badge Development, not Stable. All gen_ai content was deprecated out of the main semantic-conventions repository at v1.42.0 on 12 June 2026 and moved to a semantic-conventions-genai repository with no tagged release; the MCP attributes left in the core registry are marked deprecated for the same reason. Earlier churn is instructive: gen_ai.system became gen_ai.provider.name, prompt_tokens and completion_tokens became input_tokens and output_tokens, and gen_ai.prompt and gen_ai.completion were removed outright.
So the conclusion is narrower than "adopt OpenTelemetry". Adopt the vocabulary as your emission format, since reinventing it buys nothing an inspector values. Then pin the convention version as a configuration item under change control, normalise into your own retained schema at ingestion, and let that schema be what lives for the Chapter 4 clause 4.11 period. A stack whose field names change under you is not a records system, and a records system is what a batch retained five years after certification requires. The same separation makes an agent's actions reconstructable in regulated investigations, where multi-agent deviation triage fails unless each machine step remains traceable and attributable.
What this means in practice
Four things, in lead-time order.
Write the six thresholds before go-live, and derive them. For each signal: a metric, a subgroup breakdown, a threshold, the baseline it came from, and the action on breach. The derivation is the deliverable. A PSI threshold copied from a blog post is an unjustified acceptance criterion, and the Yurdakul and Naranjo result lets anyone show your cut-off has no stated error rate. Budget two to four weeks of a data scientist against the qualification hold-out set, and expect the abstention threshold to be re-derived once when real traffic proves less clean than the test set.
Route physical change control through the model owner. One question added to the existing GMP change control form: does this change affect data any qualified model uses as input? Name the model owner as an assessor for material, packaging, equipment and facility changes on lines where a model operates. That is a form change and an SOP revision, not a project, and it is what clause 10.1 asks for.
Instrument the human before you need the evidence. Store the model's proposal, its confidence score, the abstention flag, the reviewer's action and the interval between presentation and decision, segmented by output type. That is a schema decision at build time; retrofitting it leaves the clause 3.3 question unanswerable for every period before the retrofit.
Separate telemetry from the record on day one. Emit in the OpenTelemetry vocabulary, pin the version, and write the GMP-relevant subset into a retained store with the retention period, immutability and reason-for-change field that clause 12.2 and Chapter 4 require. A week of engineering when the system is built; a migration when it is not.
Who signs: the process subject matter expert on the sample space characterisation and the acceptance criteria, since clauses 3.1 and 4.2 assign both explicitly; QA on the monitoring procedure and every retest justification; the system owner on the configuration manifest.
What goes wrong is consistent. A threshold nobody derived, so nobody can defend it. A monitor whose bins an engineer adjusted without a change record. A supplier change that never reached the model owner. An audit trail holding the approved value but not the proposal, so the two earliest drift signals were generated and thrown away. None of these are model failures. They are operating failures, and the reason multi-agent quality workflows keep coming back to evidence and oversight is that the operating failure is the one that leaves a trace.
Questions people ask about this
- When do you have to revalidate an AI model in a GMP environment?
- No binding EU rule names a revalidation trigger for AI. The 2011 Annex 11 clause 10 requires changes to be made in a controlled manner under a defined procedure, and clause 11 requires periodic evaluation. Draft Annex 22 clause 10.1, published for consultation on 7 July 2025 and not adopted, requires every change to the model, system or process to be evaluated for retesting, with any decision not to retest fully justified.
- What is input sample space monitoring under draft Annex 22?
- Clause 10.4 of the July 2025 draft says it should be regularly monitored whether input data are still within the model sample space and intended use, and that metrics should be defined for monitoring any drift in the input data. It is separate from clause 10.3, which monitors model performance against its qualification metrics. A plan that tracks only output accuracy satisfies 10.3 and not 10.4.
- Can OpenTelemetry GenAI conventions be used as a GxP audit trail?
- They supply most of the schema and none of the record. The gen_ai and MCP attributes are all marked Development, moved to a dedicated repository at semantic-conventions v1.42.0 on 12 June 2026, and message content capture is opt-in. They carry no reason for change, no electronic signature link and no retention guarantee, so they should feed a retained GxP record store rather than being one.
- Is a rising override rate evidence of model drift?
- It is a lagging, biased signal that a quality unit should still monitor. Reviewers see inputs before labels exist, so disagreement often rises before measured accuracy falls. A falling override rate is the ambiguous case, because it can mean an improved model or a reviewer who has stopped reviewing, and only segmentation by output type separates the two.
- Do the PSI thresholds of 0.10 and 0.25 have a statistical basis?
- No. Yurdakul and Naranjo, writing in the Journal of Risk Model Validation in 2020, state that these benchmarks are used without reference to type I or type II error rates and have no support in the academic literature. The distribution of PSI depends on sample size and the number of bins, so a fixed cut-off gives a different false-alarm rate at different sample sizes.