FDA's Seven-Step AI Credibility Framework — And the Regulatory Operations AI It Never Covered
The seven steps are a good framework for an AI model that produces evidence about safety, effectiveness or quality. Two sentences in its scope section put most regulatory operations AI outside it — and running the steps anyway costs money without buying defence.
The seven steps are real, they are well built, and for a model that predicts which trial participants can skip inpatient monitoring they are exactly the right discipline. The problem is what happens when a regulatory operations team is told to apply them to an AI that classifies documents into an eCTD backbone, reconciles product registration data across forty markets, or produces a first draft of a health authority query response. That team is about to spend six months and a six-figure budget producing a credibility assessment package for a use case the guidance explicitly says it does not address.
Two sentences in section II of the draft do the work. "This guidance does not address the use of AI models (1) in drug discovery or (2) when used for operational efficiencies (e.g., internal workflows, resource allocation, drafting/writing a regulatory submission) that do not impact patient safety, drug quality, or the reliability of results from a nonclinical or clinical study." The example FDA chose for the carve-out — drafting or writing a regulatory submission — is the single highest-volume AI use case in regulatory affairs today.
That does not make the AI ungoverned. It moves it under a different and older set of instruments: the predicate rules for the records involved, the electronic submission requirements that Congress made enforceable in 2012, the quality system the company already runs, and for method, GAMP 5 Second Edition and the ISPE GAMP AI Guide. The consulting error is not caring too much about compliance. It is applying the wrong instrument, expensively, and then having nothing to show an inspector who asks a different question.
- The seven-step credibility assessment framework applies where an AI model produces information or data to support regulatory decision-making on safety, effectiveness or quality — nonclinical, clinical, postmarketing and manufacturing, and nowhere else.
- Section II excludes drug discovery and AI used for operational efficiencies that do not impact patient safety, drug quality or the reliability of nonclinical or clinical results, naming submission drafting as an example.
- Model risk is model influence multiplied by decision consequence, rated independently; FDA's own worked example downgrades influence to low because a release test independently checks the model, giving a medium overall risk from a high consequence.
- On 30 August 2026 the January 2025 text is still a draft. A Federal Register search of FDA artificial intelligence notices from 1 January 2026, run 30 August 2026, found no notice of availability for a final version.
- The binding instruments over regulatory operations AI are 21 CFR 11.1(b) and the eCTD format requirements under section 745A(a) of the FD&C Act — the second of which, unusually for a guidance, does establish legally enforceable responsibilities.
What the seven steps actually say
The framework sits in section IV.A of the draft guidance, a 23-page document issued in January 2025 by CDER with CBER, CDRH, CVM, the Oncology Center of Excellence, the Office of Combination Products and the Office of Inspections and Investigations. The Federal Register notice of availability was published on 7 January 2025 under docket FDA-2024-D-4689, with comments due by 7 April 2025.
| Step | What it asks | Where the effort lands |
|---|---|---|
| 1 | Define the question of interest the AI model will address | A single sentence, agreed with the clinical or CMC owner, not the data team |
| 2 | Define the context of use — the specific role and scope of the model | The narrower this is, the cheaper everything downstream becomes |
| 3 | Assess model risk from influence and decision consequence | The only step that sets the cost of steps 4 to 7 |
| 4 | Develop the credibility assessment plan | Model description, data description, training description, evaluation design |
| 5 | Execute the plan | The testing itself, against pre-specified acceptance criteria |
| 6 | Document results and discuss deviations from the plan | Deviations are expected and must be argued, not hidden |
| 7 | Determine the adequacy of the model for the context of use | The verdict, scoped to that context of use and no other |
The lineage matters for anyone defending the approach internally. Footnote 13 states that the concepts in steps 1 to 3 were informed by ASME V&V40, the FDA-recognised consensus standard on assessing credibility of computational modelling for medical devices. A physics-model credibility standard, borrowed for statistical models. That is why step 3 reads the way it does, and why anyone who has run a device modelling credibility argument will find the shape familiar.
The word to watch across all seven is should. The introduction says so explicitly: in Agency guidances, should means suggested or recommended, not required. Every page carries the header "Contains Nonbinding Recommendations, Draft — Not for Implementation".
Model influence and decision consequence: two ratings, set independently
Step 3 is where most internal risk assessments go wrong, because teams collapse the two axes into a single feeling about how scary the model is.
The draft defines model influence as "the contribution of the evidence derived from the AI model relative to other contributing evidence used to inform the question of interest", and decision consequence as "the significance of an adverse outcome resulting from an incorrect decision concerning the question of interest". Footnote 22 adds the constraint that most teams miss: decision consequence "should consider the question of interest, but should not consider the COU of the model". You rate the consequence of getting the decision wrong, irrespective of whether a model is involved at all.
Then the sentence that reframes the whole exercise: model risk "is the possibility that the AI model output may lead to an incorrect decision that could result in an adverse outcome, and not risk intrinsic to the model". Model card metrics, architecture novelty and parameter count are not what is being rated.
FDA's own manufacturing example is the most useful thing in the document, and it is the one to put in front of a quality director. An AI model measures fill volume in vials. Volume is a critical quality attribute, so an incorrect measurement would have a high impact on product quality and the decision consequence is high. But the manufacturer already measures fill volume on a representative sample of every batch as part of release testing. That independent test, the draft says, "would reduce the AI model influence, and therefore the model influence would be determined to be low". High consequence, low influence, medium model risk.
The engineering instruction hidden in that paragraph is worth more than the framework itself: an existing independent check downgrades model influence, and downgrading influence is the cheapest way to reduce the credibility evidence you owe. Where you cannot narrow the context of use any further, you buy risk down with a second, independent measurement rather than with more model testing. The contrast case is the clinical example in the same section, where the model is "the sole determinant" of which monitoring a participant receives — influence high, consequence high, risk high, and no amount of documentation changes that.
Which AI uses does the framework actually cover?
Section II states the positive scope in one sentence: the guidance discusses AI models in the drug product life cycle "where the specific use of the AI model is to produce information or data to support regulatory decision-making regarding safety, effectiveness, or quality for drugs". Footnote 11 removes discovery from the definition of the life cycle itself, leaving nonclinical, clinical, postmarketing and manufacturing.
Section III lists six examples of in-scope use, and they are all evidence generation: reducing animal-based pharmacokinetic, pharmacodynamic and toxicology studies; predictive modelling for clinical pharmacokinetics or exposure-response; integrating heterogeneous data sources to characterise disease presentation and progression; processing real-world or digital health technology data to develop endpoints; identifying, evaluating and processing postmarketing adverse drug experience information for reporting; and facilitating the selection of manufacturing conditions.
Note the fifth. Postmarketing adverse event identification and processing is squarely in scope, which is why the draft names the Emerging Drug Safety Technology Program as the engagement route for AI in pharmacovigilance, while warning that EDSTP "is not an avenue to seek regulatory advice on compliance with pharmacovigilance regulations". A safety team automating case intake triage is inside the framework. A regulatory operations team automating document classification is not.
Is eCTD publishing, RIM cleanup or a first-draft query response in scope?
Take the exclusion sentence apart, because the conditional at the end is doing real work.
The carve-out has two limbs. Drug discovery is excluded outright. Operational efficiencies are excluded conditionally — the exclusion holds only for uses "that do not impact patient safety, drug quality, or the reliability of results from a nonclinical or clinical study". Three named examples sit inside the carve-out: internal workflows, resource allocation, and drafting or writing a regulatory submission.
| Regulatory operations use case | Inside the carve-out? | Why |
|---|---|---|
| Document classification and eCTD granule placement | Yes | Internal workflow; the underlying study results are unchanged by where the PDF sits |
| RIM data reconciliation across registrations | Yes, normally | Data management; becomes conditional if it feeds a regulatory commitment or a variation |
| Regulatory intelligence monitoring and summarisation | Yes | No submitted information or data is produced by the model |
| Submission planning and resource allocation | Yes | Named in the carve-out verbatim |
| First-draft health authority query response | Yes, if reviewed | Drafting is named; the conditional bites if the draft's content reaches the agency unchecked |
| AI-generated analysis inside a query response | No | The model is producing information supporting a decision on safety, effectiveness or quality |
| Adverse event case triage or coding | No | Named in section III as an in-scope example |
The last two rows are the boundary, and they are not academic. A model that assembles an answer from an already-approved analysis is drafting. A model that runs the subgroup analysis the reviewer asked for, and whose output goes into the response, is producing information intended to support regulatory decision-making, whatever document it lands in. The test is not the container. The test is whether model-derived evidence reaches the reviewer's decision.
This is also where the honest limit sits. A first-draft query response generated by a language model and then materially revised by a regulatory scientist is a drafting aid. The same model's output pasted into a response with a cursory read is a route by which unreviewed model-derived content reaches a regulator, and the "do not impact" condition stops protecting you. The control that keeps the use case inside the carve-out is documented human review with a record of what changed — not a credibility assessment plan. Before you decide which side of that line a workflow sits on, it is worth costing the query response itself end to end, because the answer usually turns on how many hours of scientific review the current process already contains.
For the same reason it is a fair question whether the framework's own text anticipated any of this. It did not, in a way that is easy to check. Searching the January 2025 draft — the version served at fda.gov/media/184830/download, retrieved from the Internet Archive capture of 9 June 2026 and searched on 30 August 2026 — the terms eCTD, publishing, regulatory information management, large language model, generative and chatbot each return zero occurrences across all 23 pages. The document was written for statistical and predictive models producing evidence. It was not written about the technology most regulatory operations teams are actually deploying.
What actually binds AI in regulatory operations on 30 August 2026
Nothing in the paragraphs above says regulatory operations AI is unregulated. It says the credibility framework is the wrong instrument. Here is the right stack, with status stated for each.
| Instrument | Status on 30 August 2026 | What it reaches |
|---|---|---|
| 21 CFR 11.1(b) | Binding rule, in force since 1997 | Electronic records created, modified, maintained, archived, retrieved or transmitted under agency records requirements, and records submitted to the agency under the FD&C Act and PHS Act |
| Section 745A(a) FD&C Act plus the eCTD guidance | Binding; explicitly enforceable | The format of NDA, ANDA, certain BLA, IND and master file submissions |
| EudraLex Volume 4 Annex 11 (2011) | Binding, in operation since 30 June 2011 | Computerised systems used as part of GMP-regulated activities only |
| FDA draft AI guidance, January 2025 | Draft, non-binding; comments closed 7 April 2025 | AI producing evidence on safety, effectiveness or quality — with the section II exclusions |
| Draft Annex 11 revision and draft Annex 22 | Consultation drafts; consultation ran 7 July to 7 October 2025 | Not law; Annex 22 reaches only critical GMP applications when adopted |
| EMA–FDA guiding principles, January 2026 | Published, non-binding | Evidence generation across the drug product life cycle |
| Regulation (EU) 2026/1744 (AI Omnibus) | Binding; in force 27 July 2026 | Amends AI Act timing; high-risk Annex III to 2 December 2027, Annex I to 2 August 2028 |
Two rows deserve a second look.
The 21 CFR 11.1(b) scope is broader than most people quote it. The text reads: "This part applies to records in electronic form that are created, modified, maintained, archived, retrieved, or transmitted, under any records requirements set forth in agency regulations. This part also applies to electronic records submitted to the agency under requirements of the Federal Food, Drug, and Cosmetic Act and the Public Health Service Act, even if such records are not specifically identified in agency regulations." A submission-ready dossier is a record submitted to the agency. That is the hook, and it lands on the system holding the record, not on the model.
The eCTD requirement is the rare guidance that is legally enforceable, and it is worth knowing why. The Federal Register notice of 6 May 2015 states that "In section 745A(a) of the FD&C Act, Congress granted explicit authorization to FDA to implement the statutory electronic submission requirements by specifying the format for such submissions in guidance. Because this guidance provides such requirements under section 745A(a) of the FD&C Act, indicated by the use of the words must or required, it is not subject to the usual restrictions in FDA's good guidance practice regulations, such as the requirement that guidances not establish legally enforceable responsibilities." NDA, ANDA and certain BLA submissions had to use the FDA-supported eCTD specifications 24 months after that publication; certain IND submissions at 36 months.
So the inversion to carry into a budget meeting: the AI guidance everyone quotes is a non-binding draft that excludes your use case, and the publishing rule nobody quotes is binding and does not care about your model at all. It cares whether the output validates.
Annex 11 is worth stating precisely too, because it is routinely over-claimed. Its principle section opens: "This annex applies to all forms of computerised systems used as part of a GMP regulated activities." Submission publishing is not, in the ordinary case, a GMP-regulated activity. Annex 11 reaches your MES, LIMS and eQMS, not automatically your publishing tool, however often a vendor claims Annex 11 compliance for one.
The applicable method: GAMP 5 Second Edition, the GAMP AI Guide, and where each stops
Having removed the credibility framework, something has to fill the space. The defensible answer in 2026 is the GAMP line, used as method rather than as law.
GAMP 5 Second Edition carries Appendix D11 on AI and ML, which frames a machine-learning lifecycle from concept to operation with performance metrics and iterative training. The ISPE GAMP Guide: Artificial Intelligence followed in July 2025 at 290 pages, covering the full lifecycle of AI-enabled computerised systems, the split of responsibilities between regulated company and supplier, and ongoing monitoring and maintenance as first-class lifecycle phases rather than a post-go-live afterthought. Neither is a regulation. Both are what an inspector will recognise as current industry practice, and both scale down honestly for a low-risk system in a way a credibility assessment plan does not.
Three honest limits. The GAMP AI Guide is a paid ISPE publication, so the method your approach rests on cannot be handed to a CRO without a licence. Its lifecycle model was written with predictive and classification systems in mind, and applying it to a generative system still leaves you inventing the acceptance criteria yourself. And it carries no regulatory status: citing it shows you followed recognised practice, not that you complied with anything.
For a document classifier in a publishing pipeline that is enough. The proportionate package is a context-of-use statement, a supplier assessment, requirements risk-tagged against what a misfiled document would actually cause, a test set of real submission documents with a pre-specified accuracy threshold, a human review step with a record of overrides, and a periodic performance review. That is a GAMP 5 Second Edition risk-based package with an AI appendix, not seven steps with a credibility assessment plan and report.
Draft versus binding: the distinction that decides the budget
This is where the money is lost, so it is worth stating flatly.
On 30 August 2026 the FDA credibility framework is a draft. The header on every page says so. Comments closed on 7 April 2025. A search of Federal Register notices from the Food and Drug Administration mentioning artificial intelligence, published from 1 January 2026 onward and run on 30 August 2026, returns six documents — among them the AI-Enabled Optimization of Early-Phase Clinical Trials Pilot Program request for information under docket FDA-2026-N-4390, published 29 April 2026 with its comment period extended on 28 May 2026. None of the six is a notice of availability for a final version of FDA-2024-D-4689.
In the EU, the binding computerised systems text is Annex 11 as it came into operation on 30 June 2011. The European Commission's stakeholder consultation on a revised Chapter 4, a revised Annex 11 and a new Annex 22 on artificial intelligence opened on 7 July 2025 and closed on 7 October 2025. Neither draft is law. Annex 22, when adopted, will reach models "in critical applications with direct impact on patient safety, product quality or data integrity" — a narrower gate than most summaries suggest, and one submission publishing does not pass.
The EMA–FDA guiding principles of good AI practice in drug development, published in January 2026, are ten principles and are not binding either. Their own framing repeats the scope boundary: they are "a common set of principles to inform, enhance, and promote the use of AI for generating evidence across all phases of the drug product life cycle". Evidence generation again. Principle 2 is a risk-based approach with "proportionate validation, risk mitigation, and oversight based on the context of use and determined model risk". Principle 4 is a clear context of use. Principle 3 is adherence to relevant standards "including Good Practices (GxP)" — which is the sentence that hands regulatory operations AI back to the ordinary quality system rather than to a bespoke AI regime.
The one binding change of the last year is the AI Omnibus, Regulation (EU) 2026/1744, which entered into force on 27 July 2026 after publication in the Official Journal on 24 July 2026. It moved standalone Annex III high-risk obligations to 2 December 2027 and product-embedded Annex I obligations to 2 August 2028. It did not move prohibited practices or the Article 4 AI literacy duty, both applicable since 2 February 2025, nor Article 50 transparency from 2 August 2026. Almost no regulatory operations AI is Annex III high-risk. The literacy and transparency duties reach it anyway.
The agency draws the same line internally
There is a useful mirror on the regulator's side. FDA deployed its own internal generative AI tool, Elsa, agency-wide in June 2025, and positioned it as an aid for organisational duties rather than as part of the review determination — the operational efficiency carve-out, applied by the agency to itself. Within weeks, CNN reporting relayed by Engadget on 24 July 2025 carried accounts from FDA staff that the tool had produced non-existent studies, one describing it as hallucinating "confidently". Treat that as press reporting rather than an agency finding.
The lesson is what the carve-out does not buy. Being outside the credibility framework removes a documentation obligation. It does not remove the failure mode. A model that fabricates a citation in an internal summary is the same model that will fabricate one in a query response draft, and the only thing between that and a regulator is the human review step you designed — or did not.
What this means in practice
Write the scope decision down before you write anything else. One page: the question the AI answers, the context of use, and an explicit determination of whether the output is information or data intended to support regulatory decision-making on safety, effectiveness or quality. Cite section II of the draft, quote the exclusion sentence, state which limb applies. That page is what justifies not producing a credibility assessment plan, and it takes an afternoon. Without it, the first auditor who has read a vendor webinar asks why you skipped the seven steps and you have no answer.
Sign it at the right level. The scope determination is a regulatory judgement, not an IT one. Head of regulatory affairs or head of regulatory operations, countersigned by quality. If the determination is wrong, that is who owns it.
Budget on the right basis. A credibility assessment package for a medium-risk in-scope model is a multi-month programme with a plan, an evaluation design, a report and, if you take FDA's advice on early engagement, an agency meeting. The out-of-scope document classifier needs a validation plan, a supplier assessment, a risk-tagged requirements set, one performance qualification against a held-out set of real documents, and a periodic review. Do not let the first budget be quoted for the second job.
Design the downgrade, not the documentation. If a use case does fall in scope, the move that pays best is FDA's own vial example: find or build an independent check that reduces model influence. An existing release test, a rule-based cross-check, a mandatory human verification with a recorded decision. Each buys down evidence requirements more cheaply than more testing of the model.
Watch the three things that would change this analysis. A Federal Register notice finalising FDA-2024-D-4689, which would settle whether the exclusion survives the 2025 comments. Adoption of Annex 22, which would put a binding AI text into EU GMP for critical applications. And any FDA guidance specific to AI in postmarketing pharmacovigilance, which the January 2025 notice explicitly asked the public whether it wanted. None has happened as of 30 August 2026.
And keep the operational cases honest about their own economics. The reason regulatory operations AI attracts framework theatre is that nobody has costed the underlying process, so compliance effort becomes the only measurable thing in the business case. Working out what a query response costs today — in scientific hours, not licence fees — usually reveals that the review step you must keep for scope reasons is also the step carrying most of the cost, and that the savings are in assembly, retrieval and formatting. Which is exactly what the guidance says it does not address.
Questions people ask about this
- What are the seven steps in the FDA AI credibility assessment framework?
- Define the question of interest, define the context of use for the AI model, assess the AI model risk, develop a plan to establish credibility within the context of use, execute the plan, document the results and any deviations, then determine the adequacy of the model for the context of use. They appear in section IV.A of the January 2025 draft guidance, which remains a draft on 30 August 2026.
- Does the FDA AI guidance apply to eCTD publishing or regulatory information management?
- Not on its face. The scope section says the guidance does not address AI used for operational efficiencies such as internal workflows, resource allocation and drafting or writing a regulatory submission, where those uses do not impact patient safety, drug quality or the reliability of results from a nonclinical or clinical study. Publishing, RIM data cleanup and regulatory intelligence normally sit in that carve-out.
- What is model risk under the FDA framework?
- Model risk combines model influence, which is the contribution of evidence derived from the AI model relative to other evidence informing the question of interest, and decision consequence, which is the significance of an adverse outcome from an incorrect decision on that question. The two are rated independently. It is the risk that the output leads to a wrong decision, not risk intrinsic to the model.
- Is the FDA AI guidance legally binding?
- No. Every page of the January 2025 document is headed "Contains Nonbinding Recommendations, Draft — Not for Implementation", and it was issued under FDA good guidance practices at 21 CFR 10.115. A Federal Register search of FDA notices mentioning artificial intelligence published from 1 January 2026, run on 30 August 2026, returned no notice of availability for a final version.
- What rules do apply to AI in regulatory operations?
- The predicate rules and quality system that already applied to the process. For records submitted electronically to FDA, 21 CFR 11.1(b) and the eCTD format requirements issued under section 745A(a) of the FD&C Act. For GMP-regulated computerised systems, Annex 11 as it came into operation on 30 June 2011. For method, GAMP 5 Second Edition and the ISPE GAMP Guide: Artificial Intelligence of July 2025.