03Regulatory

Your FDA Reviewer Switched Models

A federal directive, not a validation plan, set the date on which the model behind FDA's internal assistant changed. The consequence is not that your reviewer will be misled. It is that two configurations of the same system can summarise the same dossier differently, and both summaries feed the administrative record.

On 2 March 2026 an internal FDA website banner told staff that "Gemini is already available in Elsa and will become the primary model going forward". NOTUS reported the banner the same day the HHS Office of the Chief Information Officer told employees that "as of today, users will no longer be able to log in to or access Claude through the HHS enterprise environment". Neither notice was a procurement decision. Four days earlier the President directed federal agencies to stop using Anthropic's technology, after a dispute between the Pentagon and Anthropic over autonomous weapons and domestic surveillance.

So the generative model behind the assistant your reviewer uses changed, on a deadline set by an executive directive rather than a revalidation plan. Almost everything said about it since has been alarm or reassurance, and both miss the same way. The reassurance — humans verify every step — is true and does not address the mechanism. The alarm — a hallucinating machine will misjudge your application — points at the wrong failure.

The failure mode that matters is silent. Retrieval systems do not throw errors when they degrade. Change the model that generates the answer, or the model that embeds the corpus, and the same dossier can be summarised two ways, both fluent, both plausible, neither flagged. If your submission was read partly under one configuration and partly under another, the administrative record holds two readings of one file, and nothing in it says which is which.

In short
  • The switch was politically dated, not technically scheduled: presidential directive 27 February 2026, HHS access cut 2 March 2026, internal banner naming Gemini the primary model the same day.
  • A model swap in a retrieval system is not a version upgrade. Embeddings, chunking, ranking and prompts are tuned per model, and re-tuning them changes which passages of your dossier get read at all.
  • FDA's Chief AI Officer told DIA on 20 June 2025 that Elsa "can't hallucinate" because it worked against document libraries with no internet access. Elsa 4.0, announced 6 May 2026, added secure web search.
  • The bigger change is HALO, which consolidated more than forty data sources across FDA centres, making cross-application inconsistency retrievable in one query.
  • Nothing about how you file has changed: no HALO filing route, no sponsor connectivity to Elsa, no new metadata requirement.

What happened, and when

The public record is thinner than the commentary. FDA has issued press announcements on Elsa in June 2025, December 2025 and May 2026, naming no model vendor in any of them.

DateEventSource
2 June 2025Elsa deployed agency-wideFDA press announcement
20 June 2025Chief AI Officer Jeremy Walsh tells DIA Elsa "can't hallucinate" when working against document librariesRAPS
23 July 2025CNN reports hallucinated studies and citations, quoting FDA staff anonymouslyCNN
27 February 2026Presidential directive: agencies to cease use of Anthropic technologyCNN Business
2 March 2026HHS cuts Claude access; FDA banner names Gemini the primary modelNOTUS
6 May 2026Elsa 4.0 launched; HALO consolidation completedFDA press announcement

Be precise about the sourcing, because your regulatory affairs lead will ask. The architectural description most widely quoted — a retrieval-augmented generation system built by Deloitte, evolved from an earlier CDER-GPT prototype, running in AWS GovCloud with FDA-specific document stores, embedding models and vector databases — comes from a guest column by Kimberly Chew and Michael Yang in Clinical Leader on 12 March 2026, not from FDA. The hallucination reporting is CNN quoting current and former staff anonymously; one said it "hallucinates confidently". FDA's published position is that human subject matter experts "verify all inputs, analytic processes, and output implementation". Cite the announcements as agency position and the rest as reporting.

Why swapping the model is not a version upgrade

A retrieval system has at least four model-coupled surfaces. The embedding model turns every chunk of every document into a vector; change it and the whole corpus must be re-embedded and re-indexed, because vectors from two embedding models are not comparable. Chunking is calibrated to the generator's context budget. Retrieval logic — chunk count, reranking, filters — is tuned empirically against one generator's tolerance for irrelevant context. And prompt templates are the most model-specific artefacts in the stack: instructions that suppress speculation in one model family are ignored or over-applied in another.

The Clinical Leader analysis puts the consequence plainly: developers "engineered every component of Elsa's RAG pipeline to work with Claude", and moving to a different model family "should require re-engineering and re-validating the entire system", because each model "processes information differently, which can affect how documents are retrieved, interpreted, and summarized".

Here is the part that matters for your dossier. If the embedding model changes, ranking changes. A passage in your Module 2.7.4 that ranked fourth for a reviewer's query may now rank twelfth and fall outside the retrieved set. The system reports nothing. It answers using the eleven passages it did retrieve, in complete sentences, citing real documents. The pathology is not fabrication. It is omission that reads like completeness. A summary written without the paragraph that qualifies your efficacy claim is not a hallucination; it is an accurate summary of an incomplete retrieval, and much harder to catch than an invented citation, because nothing on the page looks anomalous. It is the same reason the publishing and validation layer of a submission should stay deterministic while everything upstream of it gets smarter: non-determinism is tolerable where a human reads every output, not where the output is the record.

The guarantee that quietly expired

The most-quoted FDA reassurance is Walsh's, at the DIA Global Annual Meeting on 20 June 2025: "There are certain parts of the system the way it's designed so that when you're working on documents it forces citations. It can't hallucinate, it's not allowed to come up with figments of its imagination."

Read the conditions, not the claim. It held because Elsa worked against closed document libraries with forced citation, and because, as Walsh told the same audience, the agency did not plan to give it direct internet access. Both were engineering choices, and both have moved. The 6 May 2026 announcement lists among Elsa 4.0's new capabilities "web search through secure web access", alongside custom agents, document generation, quantitative analysis with chart creation, voice-to-text and OCR.

Once an assistant can pull external text into its context, the closed-corpus grounding that made the claim defensible no longer describes the system. That is not a criticism of the upgrade, but a warning against quoting the June 2025 assurance in 2026: the agency changed the system, not the assurance.

Does FDA hold itself to the standard it asks of you?

FDA's draft guidance on artificial intelligence supporting regulatory decision-making for drugs and biological products, issued 6 January 2025 with comments closing 7 April 2025, sets out a seven-step credibility assessment running from question of interest and context of use through model risk — model influence multiplied by decision consequence — to a documented judgement of adequacy. It remains a draft on 30 August 2026. It is guidance, it is addressed to sponsors submitting AI-derived evidence, and by its own terms it does not govern the agency's internal tooling.

So the honest formulation is not "FDA breaks its own rule". It is that the framework the agency published for models influencing regulatory decisions has no published counterpart for the model now sitting between a reviewer and forty-odd consolidated data sources: no context-of-use statement, credibility assessment plan or performance qualification for Elsa is public. The joint EMA–FDA guiding principles of good AI practice in drug development of 14 January 2026 commit both regulators to defined context of use, lifecycle management and transparency, as principles rather than obligations. You cannot compel disclosure. You can ask in writing, and a written response either documents the human validation basis or documents that none was offered.

What HALO changed that matters more than the model

Strip out the model politics and the more consequential event of 2026 is HALO — Harmonized AI & Lifecycle Operations for Data — which consolidated more than forty application and submission data sources, systems and portals across all FDA centres into one platform, queried by Elsa directly instead of requiring staff to upload documents into each chat. Walsh's framing in the announcement is the tell: "Elsa will soon become the main entrée into the FDA's systems and data."

Before HALO, an inconsistency between your current application and something you filed to a different centre three years ago surfaced only if a reviewer happened to remember it. After HALO it is a retrieval. That is a change in institutional memory, and it is permanent regardless of which model generates the prose.

Your exposure is now cross-application rather than per-application. Chew and Odette Hauke catalogue the mechanism in Clinical Leader on 26 May 2026: consolidated submission data, adverse event databases and inspection records, with no publicly disclosed role-based access boundaries between centres. They also raise an administrative-law question about Elsa 4.0's custom agents, reviewer-configured tools that can produce non-standard outputs across submissions, against the Administrative Procedure Act requirement of reasoned explanation and Motor Vehicle Mfrs. Ass'n v. State Farm, 463 U.S. 29 (1983). No litigation has tested it. Treat it as a live argument, not a precedent.

What is actually binding on 30 August 2026

Nothing in this story creates a new obligation on you, and the temptation is to let the Elsa news bleed into unrelated compliance panic.

InstrumentStatus on 30 August 2026What it reaches
EU GMP Annex 11, 2011BindingYour computerised systems in EU GMP
Draft revised Annex 11 and draft Annex 22Draft; published for consultation 7 July 2025, consultation closed 7 October 2025, not adoptedNothing yet; inspectors read them
FDA AI credibility guidance, 6 January 2025Draft; comments closed 7 April 2025Sponsor AI models supporting regulatory decisions
EMA–FDA guiding principles, 14 January 2026Published principles, non-bindingShared regulator expectations
EU AI Act, as amended by Regulation (EU) 2026/1744In force 27 July 2026; Annex III high-risk deferred to 2 December 2027, Annex I to 2 August 2028Your AI in the EU market, not FDA's

The EU AI Act does not touch Elsa: it binds providers and deployers in the EU market, not a US federal agency's internal tools. It does bind the AI you use to draft, translate and quality-check the submission, and the live obligations are the ones people forget — prohibited practices and the Article 4 AI literacy duty since 2 February 2025, GPAI obligations since 2 August 2025, Article 50 transparency since 2 August 2026. And if you are building a case that AI paid for itself in regulatory operations, the economics of health-authority response automation apply here too: a disclosed context of use, a measured baseline, an outcome someone outside the project would recognise.

What sponsors can control, and what they cannot

You cannot controlYou can control
Which model backs Elsa, or when it changes againWhether your dossier contradicts itself across modules and applications
Whether retrieval surfaces the qualifying paragraphWhether that paragraph is machine-legible: real text, tagged tables, no image-only figures
Whether an IR was AI-informed, or which prior applications were retrieved alongside this oneWhether you hold a timestamped record of every FDA interaction, and reconciled numbers across applications before filing

What this means in practice

Do not redesign your submission. FDA has announced no sponsor filing through HALO, no consolidated eCTD process tied to it, no sponsor connectivity to Elsa and no HALO-specific format or metadata. Applications still go through the Electronic Submissions Gateway or the CDRH Portal in eCTD v3.2.2 or v4.0 where eligible. Anyone selling a HALO-readiness package is selling readiness for something unannounced. Do four things instead, none wasted if the model changes again next year.

Run a contradiction pass across applications, not just within one. Reconcile the numbers that repeat — N per arm, exposure, the primary estimate and its interval, the safety denominators — everywhere they appear, including previous applications and inspection correspondence. With more than forty sources behind one query interface, this is the highest-yield hour your regulatory writer will spend. Budget it as a defined pre-publishing step, signed by the medical writer who owns Module 2.

Fix machine legibility before you fix anything clever. Scanned appendices, figures whose numbers exist only in the image, tables flattened to pictures, qualifiers surviving only in a footnote graphic: all invisible to retrieval, and OCR in Elsa 4.0 makes them semi-visible, which is worse, because a partial transcription reads as complete. This is a publishing-vendor conversation and costs almost nothing.

Keep contemporaneous records, and ask the questions in writing. Chew and Yang's checklist for the transition period of 22 April 2026 is sound and unglamorous: timestamp what you sent and what came back, request written confirmation of which AI tools processed your submission and what human validation was applied, and mark confidential commercial information explicitly rather than relying on the agency to infer it. You will often get no substantive answer; the request and the non-answer both belong in the file.

Read information requests for the shape of the query, not just the ask. An IR pairing a statement in your Module 2.5 with one from a filing to another centre, or quoting two figures back and asking which is right, is the signature of consolidated retrieval. Answer it on the merits, then treat it as a diagnostic: if one such IR arrived, the same query finds the others.

There is no defensible way to attach a number to this as a submission risk, and anyone quoting a delay probability from a model swap is inventing it. What you can tell a steering committee is narrower and true. Your dossier is now read through a retrieval system whose configuration changed under political deadline, that mediation is invisible in the record, and the only lever on your side of the wall is a dossier that says the same thing everywhere and can be read by a machine without guessing.

Questions people ask about this

Did FDA change the AI model behind Elsa?
Reporting by NOTUS on 2 March 2026 quoted an internal FDA banner stating that Gemini was already available in Elsa and would become the primary model going forward. It followed a 27 February 2026 presidential directive that federal agencies stop using Anthropic technology. FDA has not published a press announcement naming any model vendor behind Elsa, before or after the change.
Does the Elsa model switch change how I file a submission?
No. As of 30 August 2026 FDA has announced no sponsor-facing change. There is no HALO filing route, no sponsor connectivity to Elsa and no new metadata requirement. Applications still go through the Electronic Submissions Gateway in eCTD v3.2.2 or v4.0. The change is internal to the agency and affects how your dossier is read, not how it is sent.
Can FDA use AI to write an information request?
FDA states that staff verify all inputs, analytic processes and output implementation, so an information request remains a human product. What changed is what a reviewer can find quickly: Elsa 4.0 queries HALO, which consolidates more than forty application and submission data sources across FDA centres, so cross-document inconsistencies surface in one query rather than from memory.
Does the EU AI Act govern FDA's use of Elsa?
No. The AI Act binds providers and deployers in the EU market and does not reach a US federal agency's internal tools. It does reach the AI you use to prepare a submission. Prohibited practices and the Article 4 AI literacy duty have applied since 2 February 2025, GPAI obligations since 2 August 2025 and Article 50 transparency since 2 August 2026.
Is there a published validation package for Elsa?
No credibility assessment, context-of-use statement or performance qualification for Elsa has been published as of 30 August 2026. FDA's January 2025 draft guidance on AI supporting regulatory decision-making sets out a seven-step credibility framework, but it is addressed to sponsors submitting AI models, not to the agency's own internal review tooling.