01Discovery

Phase 1 Improved. Phase 2 Did Not.

Two independent analyses, two years apart, found the same shape: a large Phase 1 advantage for AI-discovered molecules and no Phase 2 advantage at all. The honest reading is that AI moved the bottleneck rather than removing it.

Two independent analyses, published two years apart and built on different cohorts, found the same shape in the data. AI-discovered molecules clear Phase 1 at a rate far above the industry norm. They clear Phase 2 at the industry norm. Nothing in between has changed.

The peer-reviewed version came first. Jayatunga and colleagues at Boston Consulting Group, writing in Drug Discovery Today in 2024, examined the clinical pipelines of AI-native biotechs and found an 80–90% Phase 1 success rate, "substantially higher than historic industry averages", against a Phase 2 rate of "∼40%, albeit on a limited sample size, comparable to historic industry averages". The IQVIA Institute's Global R&D Trends 2026 report, summarised in a 14 May 2026 post, replicated the shape on a deliberately narrower base: a vetted set of programmes with verified AI involvement rather than self-reported ones, showing a 75% Phase 1 success rate among emerging biopharma companies and Phase 2 performance "on par with their non-AI-enabled peers".

So the entire measured advantage sits at the gate that asks whether a molecule is tolerable, and none of it sits at the gate that asks whether the drug works. That is not a failure of AI. It is a precise description of what current AI does: generative chemistry optimises molecular properties, which is what Phase 1 tests. Target validity, which is what Phase 2 tests, is largely untouched.

In short
  • Two analyses converge: Phase 1 at 80–90% (BCG, peer-reviewed, 2024) and 75% (IQVIA vetted cohort, 2026); Phase 2 at roughly 40% and "on par with peers" respectively.
  • The hardest primary baseline is BIO/Informa/QLS on 12,728 phase transitions: Phase 1 52.0%, Phase 2 28.9%, Phase 3 57.8%, overall likelihood of approval from Phase 1 7.9%.
  • That same report warns Phase 1 rates "may benefit from delayed reporting or omission bias" — an alternative explanation for part of the AI Phase 1 gap that nobody prices in.
  • The one intervention with a large measured effect on Phase 2 is target genetics: relative success of 2.6x, strongest in Phases 2 and 3 and weakest in Phase 1 (Minikel et al., Nature, 17 April 2024).
  • The flagship AI clinical result, rentosertib, enrolled 71 patients and had a safety primary endpoint. Efficacy was secondary.

The 2 datasets, and what each one is measuring

BCG / Drug Discovery TodayIQVIA Institute
Published30 April 2024 (vol 29, art. 104009)Report March 2026; AI findings post 14 May 2026
CohortClinical pipelines of AI-native biotech companiesProgrammes with verified AI involvement, emerging biopharma
Phase 180–90%75%
Phase 2~40%"on par with non-AI-enabled peers"
Stated caveat"limited sample size" in Phase 2"a relatively small, validated cohort"; sample size not disclosed
Comparatorhistoric industry averagesnon-AI programmes at companies of the same segment and size

The IQVIA design is the more conservative of the two, and its result is the more useful one. It compares AI-enabled programmes with non-AI programmes at companies of the same segment and size, which removes the obvious confounder that AI-native biotechs are small, young and structurally different from large pharma. IQVIA's own framing is careful: the improvement is "visible in a segment, not yet in the whole", and industry-wide success rates did not move year over year. IQVIA also reports where the AI is actually applied — the largest share of programmes use it for molecule discovery against known targets, with smaller shares for target discovery and indication selection. That distribution is the mechanism behind the Phase 1/Phase 2 split, stated plainly by the people who built the dataset.

Which baseline are you actually comparing against?

The "40–65%" historic Phase 1 range that circulates alongside the BCG finding does not appear in the paper's abstract, which says only "historic industry averages". It shows up in secondary write-ups, including Moe Alsumidaie's 11 August 2026 piece for Clinical Trial Vanguard. Treat it as a range of published estimates, not a number.

The hardest single primary source is Clinical Development Success Rates and Contributing Factors 2011–2020, published February 2021 by BIO with Informa Pharma Intelligence and QLS Advisors, covering 12,728 phase transitions across 9,704 programmes and 1,779 companies:

TransitionSuccess raten
Phase 1 to Phase 252.0%4,414
Phase 2 to Phase 328.9%4,933
Phase 3 to NDA/BLA57.8%1,928
NDA/BLA to approval90.6%1,453
Phase 1 to approval7.9%

Two things follow. First, on this baseline AI's ~40% Phase 2 is above the 28.9% industry figure, not level with it — which tells you the BCG and BIO numbers are not measuring the same event. BIO counts a transition as failed when a programme is suspended for any reason, "including commercial viability". A small AI-native biotech with one asset and one financing runway suspends programmes for different reasons than a large sponsor rebalancing a portfolio. Anyone quoting "40% versus historic averages" as a like-for-like comparison has not read either method.

Second, and more awkward for the Phase 1 story, the BIO report attaches its own warning to exactly the number the AI cohort beats: Phase 1 rates "may benefit from delayed reporting or omission bias, as some larger companies may not deem failed Phase I programs as material and thereby not report them in the public domain." The AI-native cohort is composed almost entirely of companies for which every clinical asset is material and publicly disclosed. Some unknown fraction of an 80–90% versus 52% gap is a disclosure artefact. Neither analysis can quantify it, and neither claims to.

Why Phase 1 is the stage generative chemistry can move

Phase 1 asks a chemistry question dressed as a clinical one: is this molecule tolerated, does it reach exposure, does the pharmacokinetic profile support a dosing schedule. Those properties are computable, they have training data, and they are precisely what generative models optimise against.

The scale of that optimisation is real and worth quoting to anyone who thinks the Phase 1 result is noise. Insilico Medicine's June 2025 announcement states that its projects required synthesis and testing of only about 60–200 molecules, and that 22 candidates nominated between 2021 and 2024 went from project initiation to preclinical candidate in roughly 12–18 months against a traditional 2.5–4 years. Constraining a chemical search space to a few hundred synthesised compounds, and reaching a nomination in a year, is a genuine engineering achievement.

It is also an achievement located entirely upstream of the expensive part. IQVIA's Global R&D Trends 2026 reports industry end-to-end clinical development duration at a median of 10 years in 2025, the longest point of the past decade. Discovery compressed. The pipeline did not.

Why Phase 2 does not move

Phase 2 asks a biology question: does modulating this target change this disease in these patients. The molecule is now an instrument for testing someone's hypothesis about human pathophysiology, and the hypothesis is usually the thing that is wrong.

The attrition data has said so consistently. Harrison's analysis in Nature Reviews Drug Discovery of 218 reported failures between Phase 2 and submission over 2013–2015 attributed 52% to lack of efficacy and 24% to safety. A generative chemistry engine has no purchase on the 52%.

BenevolentAI's BEN-2293 is the cleanest illustration on record. In top-line Phase 2a results announced on 4 April 2023, the topical pan-Trk inhibitor was safe and well tolerated in mild-to-moderate atopic dermatitis and showed no statistically significant effect on either the EASI or the NRS endpoint. The molecule did what it was designed to do. The premise that pan-Trk inhibition would relieve atopic dermatitis did not hold. That is a Phase 2 failure in its purest form, and no amount of better chemistry would have changed it.

What actually moves Phase 2, and it is not chemistry

There is one intervention with a large, replicated, quantified effect on later-phase success, and it operates on target choice rather than molecule design. Minikel, Painter, Dong and Nelson, writing in Nature on 17 April 2024, estimated that drug mechanisms with human genetic support have a probability of success 2.6 times greater than those without, across 13,022 target–indication pairs that reached at least Phase 1.

The phase-by-phase breakdown is the part that matters here. In their words: "In most therapy areas, the impact of genetic evidence was most pronounced in phases II and III and least impactful in phase I, corresponding to capacity to demonstrate clinical efficacy in later development phases."

Set that against the AI profile and the two curves are mirror images. Generative chemistry helps most where genetic evidence helps least, and helps least where genetic evidence helps most. They are complementary interventions on different failure modes, and a programme that has one and not the other has bought exactly half the answer. The uncomfortable corollary for AI discovery platforms is that the measured lever on Phase 2 is a data lever — human genetics, causal gene confidence, indication–trait similarity — not a model lever.

What the flagship success actually shows

Rentosertib is the most advanced published result in the field and deserves to be read precisely rather than celebrated vaguely. The Phase 2a trial in Nature Medicine, published 3 June 2025, tested an AI-generated TNIK inhibitor against a first-in-class target identified with generative AI, in idiopathic pulmonary fibrosis, across four arms of 18, 18, 18 and 17 patients — 71 in total, in China, over 12 weeks.

The primary endpoint was the proportion of patients with at least one treatment-emergent adverse event: 72.2%, 83.3% and 83.3% across dose arms versus 70.6% on placebo. Forced vital capacity was a secondary endpoint, with a mean change of +98.4 mL (95% CI 10.9 to 185.9) in the 60 mg once-daily arm against −20.3 mL (95% CI −116.1 to 75.6) on placebo. The authors' own conclusion is that the approach "warrants further investigation in larger-scale clinical trials of longer duration".

That is a real and encouraging signal, and it is a safety-primary study of 71 patients with a nominally positive secondary endpoint. It is the strongest single data point that AI-discovered targets can work. It is not yet evidence that they work more often.

Where the pipeline stands: 117 assets, 8 through Phase 2

The cohort is still small enough to count. An analysis presented at ASCO in 2026, summarised in IntuitionLabs' 31 July 2026 pipeline review, identified 117 AI-enabled therapeutic assets across 63 companies in interventional human trials as of 1 December 2025, of which 60 (51.3%) had completed Phase 1 and 8 (6.8%) had completed Phase 2. I was unable to open the underlying conference abstract on 30 August 2026, so treat those counts as reported from a secondary summary rather than verified at source. No medicine discovered or designed with AI has been approved by any major regulator. The asset-level detail behind those counts sits in the scorecard of 117 AI-discovered assets and their disclosed clinical outcomes.

Eight completed Phase 2 readouts is not a sample from which anyone can compute a defensible success rate. Both headline analyses say so. Everyone quoting them omits it.

Reading the modelled end-to-end gain honestly

The BCG paper's modelled conclusion, widely reported from the full text, is that combining the observed AI Phase 1 and Phase 2 rates with historic Phase 3 rates lifts end-to-end probability of success from roughly 5–10% to roughly 9–18%. The article sits behind a subscription and Unpaywall records no open version, so I could not verify that pair against the full text on 30 August 2026; the abstract carries only the phase rates. Treat 9–18% as reported, not confirmed.

Take it at face value anyway and the arithmetic is unforgiving. If Phase 2 and Phase 3 rates are held at historic levels, every point of that modelled improvement is produced by the Phase 1 term. The doubling is real inside the model and it is entirely a Phase 1 doubling. In a portfolio where the largest single gate is Phase 2 at 28.9% on BIO's numbers, improving the 52.0% gate is worth having and is not where the value is locked up.

What this means in practice

If you are evaluating an AI discovery platform, a partnership, or an internal build, the diligence question is not "what is your Phase 1 rate". Every credible platform will beat the baseline there and the number is partly a disclosure artefact. Ask instead what evidence supports the target, in this indication, in humans — genetic association with a causal gene assignment, human genetic knockouts, Mendelian randomisation, tissue-level expression in the diseased state — and whether the model contributed anything to that evidence or merely designed a ligand for a target someone else chose.

Second, price the stages separately in the business case. A discovery-stage acceleration is worth its own cost line and nothing more; it is defensible to claim 12–18 months to candidate nomination and indefensible to convert that into a probability-of-success uplift downstream. The distinction between a measured saving and a modelled one is the same failure that shows up across every category of pharma AI claim, and it is worth building the same evidence standard here that applies to AI business cases a finance director will accept.

Third, get the regulatory framing right, because it will be tested. As of 30 August 2026 the FDA's "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products" remains a draft, issued 6 January 2025 with comments closed 7 April 2025; FDA guidance is non-binding even when final, and it governs the credibility of AI-generated evidence in a submission, not the biology. The joint EMA–FDA guiding principles of 14 January 2026 are principles. On the GMP side, where AI-discovery teams sometimes hear alarming second-hand claims, the binding EU text is still the 2011 Annex 11: draft Annex 22 on artificial intelligence and the draft Annex 11 revision were published for consultation on 7 July 2025, that consultation closed on 7 October 2025, and neither is law. Nothing in any of these instruments changes a probability of technical success.

Finally, be honest internally about what has been bought. AI moved the bottleneck from molecule design to target validation. That is a genuine and useful change: the constraint is now somewhere more tractable to data than it was, and the field has ten years of genetic evidence showing what a target-side intervention is worth. But a bottleneck that has moved is not a bottleneck that has gone, and the next credible claim in this field will not be another Phase 1 number.

Questions people ask about this

What is the Phase 2 success rate for AI-discovered drugs?
About 40%, on a small sample. Jayatunga and colleagues at Boston Consulting Group reported roughly 40% in Drug Discovery Today in 2024, describing it as comparable to historic industry averages and explicitly flagging the limited sample size. The IQVIA Institute reached the same conclusion in its Global R&D Trends 2026 report, finding AI-enabled programmes at emerging biopharma companies performed on par with non-AI peers in Phase 2.
Why do AI-discovered molecules do so well in Phase 1 but not Phase 2?
The two phases test different things. Phase 1 mostly tests molecular properties — tolerability, pharmacokinetics, exposure — which is exactly what generative chemistry optimises. Phase 2 tests whether modulating the chosen target changes the disease. Almost no current AI system validates that biological hypothesis, so the advantage does not carry across the gate.
Has any AI-discovered drug been approved?
No. As of 30 August 2026 no medicine discovered or designed with AI has been approved by the FDA, the EMA or any other major regulator. The most advanced published result is rentosertib, an AI-designed TNIK inhibitor for idiopathic pulmonary fibrosis, whose 71-patient Phase 2a trial appeared in Nature Medicine on 3 June 2025 with a safety primary endpoint.
What actually improves Phase 2 success rates?
Human genetic evidence for the target. Minikel and colleagues, in Nature on 17 April 2024, estimated that drug mechanisms with genetic support have a 2.6 times greater probability of success, and found the effect was most pronounced in Phases 2 and 3 and least impactful in Phase 1 — the exact inverse of the AI profile.
Does the FDA have binding rules on AI in drug development?
Not yet. The FDA guidance "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products" was issued as a draft on 6 January 2025 and the comment period closed on 7 April 2025. As of 30 August 2026 it remains draft, and FDA guidance is non-binding in any case. The joint EMA-FDA guiding principles published on 14 January 2026 are principles, not requirements.