Digital Twins, External Controls and What EMA Actually Qualified
CHMP qualified a covariate adjustment method for randomised trials with continuous outcomes. It explicitly declined to qualify the model that produces the score, and no regulator has yet accepted a machine-generated control arm in place of a real one.
Ask a clinical development lead what EMA qualified in 2022 and you will usually hear some version of "digital twins, so you need fewer patients on placebo". The second half is roughly right. The first half is wrong in a way that matters when a health authority asks you to defend the design.
The CHMP qualification opinion for Prognostic Covariate Adjustment (PROCOVA), adopted 15 September 2022 after a public consultation that ran 22 March to 3 May 2022, qualifies a statistical method: use a model trained on historical control data to predict each randomised participant's outcome under control, then adjust for that single predicted value as a covariate in an ANCOVA. The context of use is narrow and stated on page 2. It is randomised controlled trials with continuous outcomes. Everyone is still randomised. There is still a control arm. What changes is the residual variance, and therefore the number of participants you need.
The opinion is also explicit about what it does not cover. CHMP wrote that it "cannot qualify a formalised procedure for prognostic model development in Step 1 as part of the PROCOVA method", and recorded that the applicant had put model development "explicitly out of scope of this qualification procedure". The machine learning is the part nobody qualified. Non-linear analysis models and treatment-by-covariate interactions are out of scope too. So is the thing most people think happened: replacing control participants with generated ones.
- CHMP qualified PROCOVA on 15 September 2022 for randomised trials with continuous outcomes only; binary, count and time-to-event endpoints were outside the context of use.
- The qualification covers the adjustment, not the model: CHMP declined to qualify prognostic model development, which the applicant had placed out of scope.
- The opinion's own worked example cut a placebo arm from 164 to 137 and total sample size from 402 to 343 — a 15 per cent reduction, not the larger figures quoted in vendor material.
- FDA declined to accept a PROCOVA Letter of Intent into ISTAND, the qualification programme made permanent on 31 July 2025, on the basis that existing covariate guidance already covers it.
- ICH E6(R3) Annex 2, effective in the EU on 15 January 2027, puts a seven-item fitness-for-purpose test around real-world control data — including the validation status of the tools that collected it.
What CHMP actually qualified
The operative sentence in the opinion is careful. CHMP qualifies PROCOVA "as prognostic score adjustment", and says the procedures "could enable increases in power or precision of treatment effect estimates in controlled randomised clinical trials with continuous outcomes". On sample size the wording is conditional rather than permissive: the reduction in residual variance "may in principle be taken into account to reduce sample size, if it can be ensured that the calculation is considering uncertainties in the assumptions made, and if the resulting sample size is large enough to meet other relevant purposes of the clinical trial apart from the primary hypothesis test".
That second condition is the one that gets forgotten. A phase 3 trial is not only a hypothesis test. It is also the safety database, the exposure denominator, the subgroup evidence and, in Europe, part of the material an HTA body will read. Shrink to the statistical minimum and you can win the primary endpoint while losing the label.
CHMP was equally clear that it was not anointing a product. It "does not intend to single out any specific method for statistical modelling using adjustment for covariates as 'the' method to be used". The opinion tells the trial statistician to compare three paths — no adjustment, ANCOVA with pre-specified covariates, or PROCOVA — and says an advantage over plain ANCOVA "should be justified", which in practice means the prognostic score has to capture a non-linear relationship that a linear combination of your usual baseline covariates would miss.
One genuine reassurance sits in the opinion and deserves repeating, because it is the reason this method is defensible at all: "Type I error control, unbiased effect estimation and confidence interval coverage are not dependent on the choice or performance of the prognostic model." A bad model costs you efficiency. It does not manufacture a false positive. That property is what separates prognostic covariate adjustment from a synthetic control arm, where a bad model does exactly that.
The number that is not in the opinion
Search the qualification opinion for a 35 per cent control arm reduction and you will not find it. What you find is the applicant's own worked example, a re-analysis of a published Alzheimer trial, laid out in Table 5.
| Design | Active | Placebo | Total | Reduction |
|---|---|---|---|---|
| Unadjusted | 238 | 164 | 402 | — |
| Random forest prognostic score | 217 | 144 | 361 | 10% |
| Deep learning prognostic score | 206 | 137 | 343 | 15% |
Source: CHMP qualification opinion for PROCOVA, Table 5. The placebo arm falls by 27 participants. The opinion also notes that in this design the confidence intervals for the effect on the CDR endpoint were 6 per cent wider, because sample size had been estimated from performance on a different endpoint.
The larger figures come from elsewhere. The frequently quoted "up to 35 per cent smaller control arm" traces to a company retrospective analysis of a phase 2 crenezumab study, published as vendor communication rather than in the qualification dossier. A 2025 conference abstract in Alzheimer's & Dementia, Kusiak and colleagues on a simulated TRAILBLAZER-ALZ 2 trial, reports up to 24 per cent fewer control participants and up to 15 per cent more power on secondary endpoints — from a synthetic dataset constructed to mirror the real cohort, authored by the model vendor. None of that is fraudulent. It is simply a different evidentiary tier from a CHMP opinion, and if you put the 35 per cent figure in a design justification you should expect to be asked which document it came from.
The honest planning range is the one the opinion gives as a rule of thumb: variance falls by roughly one minus the square of the correlation between prognostic score and outcome. A score correlating at 0.5 buys about a 25 per cent variance reduction; at 0.8, about 64 per cent. Whether your disease area supports a score at 0.8 is an empirical question, and CHMP declined to answer it, saying it "cannot issue a statement about the precision of prognostic models in general and over therapeutic areas".
Does FDA have an equivalent?
Yes and no, and the detail is worth knowing because it is often stated wrongly in both directions.
FDA does run a qualification programme for novel drug development tools that do not fit the biomarker, clinical outcome assessment or animal model categories: the Innovative Science and Technology Approaches for New Drugs programme, launched in November 2020 as a pilot and made a permanent qualification programme on 31 July 2025. FDA has said it has accepted eight submissions, including three AI-based tools and one novel statistical methodology.
PROCOVA is not among them. By the applicant's own published account of CDER's response, FDA declined to accept the Letter of Intent into ISTAND, on the reasoning that "CDER's current feedback is that PROCOVA does not appear to deviate from our guidance" — that is, the method was judged to be already inside the four corners of the May 2023 final guidance on adjusting for covariates in randomised clinical trials, which contemplates adjustment for a composite prognostic index constructed from previous studies. Not qualifying something because it is unremarkable is not the same as qualifying it, and it is a long way from endorsing the model that builds the index.
The vendor's own position is worth quoting back to anyone who oversells this internally. Its regulatory strategy note says digital twins augment rather than replace control arms; the prognostic score is a covariate inside a randomised design, not a substitute participant.
For the harder question — a control group made of other people's data — FDA's position remains a draft. Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products was issued as draft guidance in February 2023 and, as of 30 August 2026, has not been finalised. Draft guidance represents FDA's current thinking and is not binding on the agency or on you.
Where Europe actually is on external controls
Three separate instruments, three different statuses, and mixing them up is how sponsors lose credibility in a scientific advice meeting.
| Instrument | Status on 30 August 2026 | What it governs |
|---|---|---|
| CHMP qualification opinion, PROCOVA | Adopted 15 September 2022 | Covariate adjustment in randomised trials, continuous outcomes |
| Reflection paper on single-arm trials as pivotal evidence | Published, in force as EMA guidance | Justifying the absence of randomisation |
| Reflection paper on external controls | Concept paper only; consultation ran 25 July – 31 October 2025 | Not yet published |
| ICH E6(R3) Annex 2 | ICH Step 4 on 3 June 2026, CHMP adopted 25 June 2026 | Effective in the EU 15 January 2027 |
| Virtual control groups in rat dose-range finding | Draft qualification opinion, 31 March 2026, comments to 12 May 2026 | Preclinical only |
That last row is the sharpest illustration of the gap. The one place a European regulator has publicly entertained replacing a control group outright is non-GLP rat dose-range finding studies, and even that was issued as a draft qualification opinion for consultation, not a final one. Nothing comparable exists for human efficacy.
Meanwhile the concept paper for a reflection paper on external controls closed its consultation on 31 October 2025 and the reflection paper itself has not appeared. If your 2027 filing strategy assumes an EMA external-control guideline, it is assuming a document that does not exist.
What Annex 2 section 3.5.1 requires
This is the part that changes operational reality, and it is now settled text rather than draft. ICH E6(R3) Annex 2 reached Step 4 on 3 June 2026; EMA records CHMP adoption on 25 June 2026 and entry into effect on 15 January 2027. Section 3.5.1(a) defines fitness for purpose as reliability — accuracy, completeness, provenance, traceability — plus relevance, meaning the key data elements exist at all. Section 3.5.1(b) then lists seven considerations that "may inform the determination of whether a given RWD source is fit for purpose".
| Item | What it asks |
|---|---|
| (i) | Variability of formats, terminologies and standards across sources |
| (ii) | Absence of standardised timing of assessments — clinical practice timing is driven by the patient's condition |
| (iii) | Comparability between trial data and the control group using RWD: baseline factors, selection bias, consistency of clinical evaluation |
| (iv) | Missing data and intercurrent events that are hard to ascertain, cross-referenced to ICH E9(R1) |
| (v) | Overall quality of data from clinical practice: database structure, vocabularies, coding |
| (vi) | De-identification methods used |
| (vii) | Fitness for purpose of the systems and tools used for collection and acquisition of RWD, including validation status as appropriate |
Item (iii) is where the external control arm lives. Item (vii) is the one that catches teams by surprise: the pipeline is in scope, not just the data. If a registry front end, a digital health technology or an extraction tool sits between the patient and your analysis dataset, its validation status is part of the fitness-for-purpose argument. For sponsors using large language models to normalise or curate that data, the burden is the same, and the failure modes of LLM-curated real-world evidence show up as biased estimates rather than obvious errors.
Two more Annex 2 provisions bind harder than most sponsors expect. Section 3.4.1 requires that arrangements with data owners "allow regulatory authorities to access individual-level data and source records to support regulatory inspections" — an inspection clause you must negotiate into the registry or EHR agreement before the trial, not after. And for higher-criticality uses, "simply assessing the fitness for purpose of the RWD source to be utilised may be insufficient", so sponsors may need to confirm "whether a clinical event occurred, was assessed, or was documented" against source records.
One thing Annex 2 does not do is regulate the model. Searching the ICH Step 4 final text published 3 June 2026, on 30 August 2026, the phrases "artificial intelligence", "machine learning", "algorithm", "digital twin", "synthetic" and "prognostic" return zero occurrences; "external control" appears once. Model governance for a GCP system still comes from ICH E6(R3) Annex 1 section 4 and from EMA's Guideline on computerised systems and electronic data in clinical trials (EMA/INS/GCP/112288/2023), adopted by the GCP Inspectors Working Group on 7 March 2023 and in effect six months after its 9 March 2023 publication.
What an external control arm costs in patients you cannot use
The strongest available evidence on this is not a vendor deck. It is an EMA-hosted overview of external controls in submissions to the agency, presented on 11 November 2025 by a CHMP and SAWP alternate member, which walks through the Abecma dossier.
The retrospective study built to provide the external comparison started from 1,949 patients meeting the prior-therapy requirements. Of those, 528 received a new therapy. 190 met the key criteria, including at least one disease assessment. The matched cohort was 76 to 80 patients. More than 90 different regimens appeared among the 190 eligible patients.
CHMP accepted the comparison and then said what it thought of it: the comparisons were "limited by several factors including the rather long time period (up to 60 days from the index date) allowed for the collection of baseline data, the overlapping recruitment periods... the large proportion of missing data (up to 30%) for some included co-variates and several co-variates excluded from the PS model due to >30% missing data". The conclusion was that "the true magnitude of the treatment effect... cannot be reliably ascertained", but that the efficacy signal was large enough to carry the benefit-risk assessment anyway.
That is the actual regulatory bargain for external controls, and it holds across the Yescarta and Zolgensma precedents in the same presentation: they work when the effect size is enormous, the natural history is grim and well characterised, and randomisation is genuinely hard to justify. The same presentation is blunt about the other case — "attempts to rescue negative trials based on comparisons to external data are generally not accepted".
What this means in practice
If you are designing a trial in the next six months, the sequence is not complicated.
Decide first whether you are doing covariate adjustment or an external control. They are different regulatory objects with different burdens, and the marketing language blurs them. Covariate adjustment inside a randomised trial is ordinary statistics with a qualification opinion behind it; an external control is a request to accept non-randomised evidence and needs a justification for why randomisation was not done.
If it is covariate adjustment, pre-specify it in the protocol and the statistical analysis plan before unblinding, and pre-specify the prognostic model as a frozen artefact with a version identifier. Because CHMP declined to qualify model development, the model is your evidence to produce: training data provenance, the independence of the training set from the trial, measured correlation on held-out data, the deflation factor applied to guard against over-optimism, and the sensitivity analysis showing what happens if the correlation is lower in your trial than in development. Budget for the case where it is lower — CHMP warns that "outcomes from historical data may not allow prediction of control arm outcomes of future trials in case of changes in the therapeutic landscape", which in oncology and immunology means any shift in standard of care since the historical data were collected.
Check your endpoint. The qualification covers continuous outcomes. If your primary endpoint is binary, a count or time-to-event, you are outside the context of use and should say so in your scientific advice briefing package rather than let an assessor find it.
If it is an external control, the Annex 2 clock is the constraint. Trials starting after 15 January 2027 in the EU face the seven-item assessment and the inspection-access requirement, and data access agreements take months to renegotiate. Start with item (vii): list every system between the patient and the analysis dataset and record its validation status. Most sponsors can answer that for the EDC and cannot answer it for the registry extract.
Sign-off is a three-signature problem in most organisations — the trial statistician owns the adjustment method and the sample size justification, quality assurance owns the fitness-for-purpose file and the validation status of the acquisition tooling, and the regulatory lead owns the story about why randomisation was reduced or absent. If those three have not read the same version of the Annex 2 text, you will find out during the inspection rather than during the design.
And the sentence to keep in the room: on 30 August 2026, no medicines regulator has qualified a machine-generated control arm as a substitute for randomised control participants in a pivotal efficacy trial. What exists is a qualified way to need fewer of them, and a very old, very demanding route for using other people's data when randomisation is not possible.
Questions people ask about this
- Did EMA approve digital twins to replace clinical trial control arms?
- No. The CHMP qualification opinion for PROCOVA, adopted 15 September 2022, covers prognostic covariate adjustment in randomised controlled trials with continuous outcomes. Every participant is still randomised and every trial still has a real control arm. The opinion allows the prognostic score to be counted when estimating sample size, so the control arm can be smaller, not absent.
- Does FDA have an equivalent of the EMA qualification opinion for PROCOVA?
- FDA runs the ISTAND qualification programme, made permanent on 31 July 2025. Unlearn submitted a Letter of Intent for PROCOVA and, by the company account of CDER feedback, FDA declined to accept it into ISTAND on the grounds that covariate adjustment is already covered by the May 2023 final guidance on adjusting for covariates. There is no FDA qualification of PROCOVA.
- What does ICH E6(R3) Annex 2 say about external control arms?
- Annex 2 reached ICH Step 4 on 3 June 2026, was adopted by CHMP on 25 June 2026 and comes into effect in the EU on 15 January 2027. Section 3.5.1(b) lists seven considerations for deciding whether a real-world data source is fit for purpose. Item (iii) covers comparability between trial data and a control group drawn from real-world data; item (vii) covers the validation status of the collection tools.
- How much can a prognostic score actually shrink a control arm?
- In the worked example inside the CHMP opinion itself, a re-analysis of an Alzheimer trial, the placebo arm fell from 164 to 137 participants and total sample size from 402 to 343, a 15 per cent reduction. Larger figures circulating in industry come from company retrospective analyses and simulations, not from the qualification opinion.
- Is there an EMA guideline on external control arms yet?
- Not as of 30 August 2026. EMA published a concept paper for a reflection paper on external controls and ran a consultation from 25 July to 31 October 2025. The reflection paper itself has not been published. Until it is, sponsors work from the existing reflection paper on single-arm trials and from case-by-case scientific advice.