02Clinical

Every Disclosed AI Cycle-Time Number in Clinical Development, and What Each Measures

Four numbers carry almost every AI business case in clinical development. All four time the same thing — how long it takes to produce a first draft — and the largest of them came from retrieving pre-approved sentences, not from a smarter model.

Four numbers do almost all the work in AI business cases for clinical development: Novo Nordisk's 90% cut in clinical study report writing time, the 97% first-draft reduction reported for AutoIND with Takeda, AstraZeneca's 85% on protocol authoring, and Novartis's 83–87% on compliant protocol generation. Each is real, each is attributable to a named organisation, and each measures the same narrow thing — the time to produce a first draft, with a qualified human still owning review and approval.

None of them measures time to regulatory approval. Two of the four say so explicitly. And the largest and most cited of them, Novo Nordisk's 90%, did not come from a more capable model. It came from retrieving domain-expert-approved text blocks and slotting case-specific variables into them — constrained generation over a controlled library, with vector similarity picking which pre-approved snippet matches which statistical output.

That distinction decides whether your business case survives contact with a finance director. A drafting-time saving converts to a cycle-time saving only when drafting sat on the critical path, and it converts to a revenue number only when the whole downstream chain — review, QC, sign-off, publishing, agency clock — moves with it.

In short
  • All four headline numbers — Novo Nordisk 90%, AutoIND 97%, AstraZeneca 85%, Novartis 83–87% — time first-draft production, not approval.
  • Novo Nordisk's CSR system works by vector-matching pre-approved text snippets to statistical outputs with full source lineage; the residual effort moved into QA, not out of the process.
  • The AutoIND 97% has a measured numerator and an estimated denominator: the ~100-hour manual baseline came from Takeda writers' experience, not a timed control arm.
  • The same study scored AI drafts at 69.6% and 77.9% against submission-ready quality, with 3–5x verbosity inflation and 35% of documents missing essential study design elements.
  • Searched on 30 August 2026, I found no sponsor disclosure of a regulatory review or approval time reduced by AI — and the AutoIND authors state they did not measure it.

The 4 numbers, and what each one timed

DeploymentHeadline numberWhat was actually timedBaselineHuman review
Novo Nordisk, NovoScribe90% cut in CSR writing time; ~12 weeks → ~10 minutesGeneration of a CSR draft from statistical outputsInternal historic compile time; staff writers averaged 2.3 CSRs/yearRetained; "the rest of the time is spent in QA"
Weave Platform AutoIND with Takeda97% cut in first-draft compositionDrafting of eCTD modules 2.6.2, 2.6.4, 2.6.6 for 2 INDs~100 hours, estimated by Takeda writers, not measuredRequired; drafts scored 69.6% / 77.9% of submission-ready
AstraZeneca, intelligent protocol tool85% reduction in document authoring time "in some cases"Protocol document authoringNot disclosedRetained; outputs framed as first drafts
Novartis with AWS83–87% acceleration in generating compliant protocolsProtocol generationNot disclosedNot disclosed in the available material

Two things are visible in that table before any of the detail. First, every "what was timed" column says drafting. Second, the quality of the sourcing falls off sharply from left to right. The Novo Nordisk figure appears in two vendor-published customer stories with named executives and consistent numbers. The AutoIND figure comes with a full methods section, a scoring rubric and a limitations section. AstraZeneca's 85% carries the qualifier "in some cases" and no published method. The Novartis figures come from a recorded conference session with AWS, summarised in a third-party database, with no primary Novartis publication and no stated measurement protocol — quote them, if at all, as a directional claim, not a benchmark.

What Novo Nordisk's 90% actually measures

The mechanism matters more than the number. Per the MongoDB customer case study, NovoScribe "generates validated text based on defined content rules and statistical output", and "Atlas Vector Search calculates the similarity of each text snippet to the relevant statistics. This combined with the LLM output draft the CSR." The Anthropic customer story describes the same architecture as retrieval-augmented generation over domain-expert-approved text with case-specific variables, running on Amazon Bedrock.

Read that carefully. The system is not writing a CSR the way a medical writer writes one. It is selecting from a library of sentences that domain experts have already approved, matching them to the statistical outputs of a specific trial, and binding variables into them. The intelligence is in the library and the retrieval, not in the generation. That is why the output is defensible under GCP: every claim traces to an approved source, and Novo reports that "full lineage of all sources are presented, enabling the authors to verify accuracy."

The rest of the disclosed detail sharpens the picture rather than softening it. The programme started in mid-2023 and is maintained by an 11-person team. Device verification protocols showed a 95% resource reduction. At the time of the MongoDB write-up the system covered around 30% of all Novo Nordisk CSRs with an expectation of passing 90%. The baseline is the most revealing figure of all: staff writers averaged 2.3 CSRs per year. That is the number a cycle-time business case should anchor on, because it describes the constraint the tool relieves — writer capacity — rather than the elapsed time of a trial.

And the honest quote from Waheed Jowiya, Novo Nordisk's digitalisation strategy lead, is the one every deck leaves off: Claude "helped us cut writing times on CSRs by 90% so we can get documentation directly into human hands for review and approval." The tool ends where the regulated act begins.

The AutoIND study is the only one that published its method

The Weave Platform and Takeda preprint on human-AI collaboration in regulatory writing, posted to arXiv in September 2025, is the most useful document in this entire evidence base, because it is the only one that shows its working — and what it shows is a headline number with a soft denominator.

The measured side is solid. AutoIND produced first drafts of the IND nonclinical written summaries (eCTD modules 2.6.2, 2.6.4 and 2.6.6) in 3.7 hours from 18,870 pages across 61 source documents for the first IND, and 2.6 hours from 11,425 pages across 58 documents for the second. Those times were recorded directly.

The comparator was not. In the authors' own words, manual drafting times "were estimated based on the experience of regulatory writers (≥ 6 years' experience) at a multinational biotechnology company (Takeda Biopharmaceuticals)", and "manual times were used only as industry-standard benchmarks to contextualize efficiency." A 97% reduction computed against a recalled benchmark is not the same class of evidence as a randomised time-and-motion study, and the paper does not pretend otherwise.

The quality findings are where the article's argument is settled. A single blinded assessor scored the drafts across seven categories on a 0–3 scale, yielding 69.6% and 77.9% against submission-ready expectations, with no critical regulatory errors detected. But the expert analysis behind those scores found a consistent failure pattern: 3–5x word-count inflation over human-written equivalents, 35% of documents missing essential study design elements, a 32% failure rate in maintaining logical information flow, and a 100% failure rate on specific elements such as dose formulation analysis in GLP studies. Those defects transfer work to the reviewer. If you want the full breakdown of how the speed number and the quality number sit against each other, that is the subject of a separate reading of the AutoIND quality scores.

The limitations section then closes the loop on cycle time in one sentence: "The study did not assess long-term regulatory outcomes, such as FDA review comments or approval timelines, which would provide ultimate validation of document quality." The authors also note the study covered two INDs from a single therapeutic area, modality and company, evaluated by a single assessor.

Has any sponsor disclosed a shorter time to approval?

Not that I can find. Searching on 30 August 2026 across sponsor communications, earnings coverage and trade press for a disclosed reduction in regulatory review time, filing-to-approval time or agency question volume attributable to AI, I found no such disclosure from any sponsor. This is an absence-of-evidence finding, not proof that none exists — but the absence is consistent, and two of the four programmes here say explicitly that they did not measure it.

There is a structural reason. Review clocks are set by statute and user-fee goal letters, not by how fast the dossier was drafted. Compressing authoring moves the submission date forward; it does not shorten the assessment that follows. The only mechanisms by which AI could shorten the regulatory phase are better first-cycle quality — fewer information requests, fewer major objections — and that is precisely the outcome nobody has published, because it requires a multi-year cohort and a credible control.

The regulator's own side of the desk has no published cycle-time number either. FDA launched its internal generative AI assistant, Elsa, in June 2025, and trade coverage of its accuracy problems reports reviewers still flagging fabricated citations after later releases. Separately, on 29 April 2026 FDA announced steps to implement real-time review of clinical trial data, with an AstraZeneca-sponsored phase 2 mantle cell lymphoma trial and an Amgen-sponsored phase 1b small cell lung carcinoma trial as first participants. That programme is about data plumbing and concurrent signal reporting, not about AI drafting, and it has produced no cycle-time figure yet.

What the regulators have actually put in writing

The instrument-by-instrument position on 30 August 2026, because vendors routinely quote drafts as though they bind:

InstrumentStatus on 30 August 2026Relevance here
EMA guideline on computerised systems and electronic data in clinical trialsBinding practice; effective 9 September 2023The operative validation text for a GCP system that drafts CSRs or protocols
ICH E6(R3) Principles and Annex 1Effective in the EU 23 July 2025Risk-proportionate quality; sponsor oversight of systems and vendors
ICH E6(R3) Annex 2ICH Step 4 on 3 June 2026, CHMP adoption 25 June 2026, effective in the EU 15 January 2027Not yet in effect; plan for it, do not cite it as current
FDA draft guidance on AI to support regulatory decision-makingDraft, published January 2025; comment period closed 7 April 2025; still draft in trade coverage dated 28 July 2026Seven-step credibility assessment framed around context of use
FDA/EMA Guiding Principles of Good AI Practice in Drug DevelopmentPublished 14 January 2026; 10 principles; not bindingShared vocabulary for risk-based justification
EU GMP Annex 11 (2011)BindingGMP scope, not GCP; the revision published for consultation on 7 July 2025 is not law
Draft EU GMP Annex 22 (AI)Consultation draft; published 7 July 2025, consultation closed 7 October 2025Not law; inspectors read it anyway
EU AI Act, as amended by Regulation (EU) 2026/1744In force 27 July 2026Annex III high-risk obligations moved to 2 December 2027; Annex I to 2 August 2028

Two practical consequences. First, an AI system that drafts a CSR at a drug sponsor is a GCP computerised system, so the framework is the EMA 2023 guideline plus GAMP 5 Second Edition and the ISPE GAMP AI Guide of July 2025 — not Annex 22, which sits in GMP scope and is not adopted, and not Computer Software Assurance, which is US medical-device scope under 21 CFR Part 820 and does not apply to a drug sponsor's GCP systems.

Second, the EU AI Act deferrals matter less to this use case than the coverage implies. The high-risk timetable slipped, but a drafting assistant whose output is reviewed and approved by a qualified writer is unlikely to be an Annex III system in the first place. What already applies is the Article 5 prohibited-practices regime, in force since 2 February 2025, the Article 4 AI literacy duty, and the GPAI provider obligations that have applied since 2 August 2025.

The review rate decides the business case, not the drafting number

Here is the arithmetic that vendor decks skip. Novo Nordisk's own description is that the model takes minutes and "the rest of the time is spent in QA". So the process time did not go to zero; it changed shape. The AutoIND evidence points the same way with the opposite sign: a draft that is 3–5 times too long and missing 35% of essential design elements costs a reviewer more, not less, per page.

That means the sensitivity in any AI clinical development ROI model sits in one variable — how much human verification each output needs — and not in the drafting reduction that the headline advertises. A 90% drafting cut with 100% line-by-line verification of an inflated draft can net out to a smaller saving than a 50% drafting cut against a tight, lineage-traceable output that a reviewer can spot-check. Novo Nordisk's design choice makes sense in exactly this light: constraining generation to approved text blocks is a way of buying down the review cost, not a way of writing faster.

The value side is easy to over-claim. Novo Nordisk's own figure, as reported in the MongoDB study, is that each day sooner a medicine reaches market is worth around $15 million to the company. Tufts CSDD data reported in trade analysis puts the mean direct cost of conducting a trial at roughly $40,000 per day, and $55,716 per day for phase 3. Those numbers are only reachable if the document you accelerated was the thing holding up the milestone. For a CSR that is finalised months after last-patient-last-visit and sits inside a submission assembled over quarters, it usually is not. Bank the capacity saving, which is real and measurable, and stop selling the launch-date saving, which nobody has evidenced.

What is the real pilot-to-production rate?

There is no trustworthy pharma-specific figure, and the one most often quoted does not survive checking. The widely repeated claim that around 24% of pharma proofs of concept reach production could not be traced to the source it is usually attributed to when I checked it against that source; treat it as folklore until someone produces the underlying survey.

What can be sourced is broader. Deloitte's State of AI in the Enterprise 2026, covering 3,235 business and IT leaders across 24 countries and six industries including life sciences, found that only 25% of organisations had moved 40% or more of their pilots into production, with more than half expecting to cross that line within months, and only 21% reporting a mature agent-governance model. MIT's NANDA analysis, reported in mid-2025, found 95% of enterprise generative AI pilots delivered no measurable P&L impact. Benchling's 2026 biotech survey reported 40% of AI pilots reaching scaled deployment and 17% able to demonstrate measurable value in discovery.

The useful reading is not that pilots fail. It is that the four deployments in this article all share the property the failures lack: a unit of measurement (per CSR, per IND module, per protocol), a named baseline, and an explicit human sign-off point.

What this means in practice

Instrument the baseline before you build anything. If you cannot state today's median hours per CSR section, per protocol, or per module 2.6.4 — measured, from your own document management audit trails, not recalled by writers — you will end up with a number as soft as the ~100-hour comparator in the AutoIND paper, and your QA director will say so in the first review meeting.

Define the measured unit and the boundary in the charter. "Time from statistical outputs available to first complete draft in the authoring system, excluding review cycles" is a defensible claim. "Cycle time reduction" is not, unless you can name the milestone that moved.

Put your engineering effort on the review side. The Novo Nordisk architecture is the model to copy: a curated library of approved text, retrieval that binds each generated passage to the statistical output it describes, and lineage presented to the author so verification is a check rather than a re-read. That is what makes a sampled review defensible instead of a 100% line-by-line one, and the review rate is the variable your business case actually turns on.

Get the validation stack right in the first slide, because the wrong framework ends the conversation. For a GCP authoring system: EMA computerised systems guideline effective 9 September 2023, GAMP 5 Second Edition, ISPE GAMP AI Guide of July 2025, with a documented context-of-use statement and pre-defined acceptance criteria on an independent test set. ICH E6(R3) Annex 2 comes into effect on 15 January 2027 and belongs in the roadmap, not in the current-state assessment. The human approver stays the release authority, and the evidence that this is not a formality is in the seven-category quality scoring behind the AutoIND speed claim: 69.6% and 77.9% against submission-ready, with no critical regulatory errors but pervasive structural defects.

Decide who owns the number before you publish it internally. The four disclosures in this article were made by digitalisation and R&D IT functions, not by regulatory affairs, which is why they are all authoring metrics. If your steering committee wants a submission-date claim, the person who has to sign it is the head of regulatory operations, and they will want to see the milestone that moved and the assumptions behind it — not a percentage lifted from a vendor's customer story.

Expect the failure mode to be verbosity, not hallucination. Every published quality assessment here found the model producing more text than a human would, with structural drift and omitted design elements, rather than inventing facts. Budget reviewer time for cutting, and write an acceptance criterion on output length before your first pilot document reaches a medical writer.

Questions people ask about this

What did Novo Nordisk actually reduce by 90%?
Writing time for clinical study reports. NovoScribe generates a CSR draft in around ten minutes against a previous compile time of roughly twelve weeks, but Novo Nordisk describes the output as documentation put "into human hands for review and approval". The remaining effort moved into quality assurance. No reduction in regulatory review or approval time has been claimed.
Has any pharma company published a shorter time to approval because of AI?
Not as of 30 August 2026, on the evidence I could find. Every disclosed figure in clinical development measures authoring or drafting time. The AutoIND study with Takeda states explicitly that it did not assess FDA review comments or approval timelines. Review clocks are set by statute and user-fee goals, not by how quickly a dossier was drafted.
How was the AutoIND 97% figure calculated?
AutoIND drafting times for eCTD modules 2.6.2, 2.6.4 and 2.6.6 were measured directly: 3.7 hours and 2.6 hours for two INDs. The roughly 100-hour manual comparator was not measured. It was estimated from the experience of Takeda regulatory writers with six or more years in the role, and the authors describe it as an industry-standard benchmark used to contextualise efficiency.
Does the EU AI Act cover an AI tool that drafts clinical study reports?
Rarely through the high-risk regime. Regulation (EU) 2026/1744, in force 27 July 2026, moved Annex III high-risk obligations to 2 December 2027 and Annex I to 2 August 2028. A drafting assistant with a qualified human approver is unlikely to sit in Annex III at all. Article 5 prohibitions, the Article 4 AI literacy duty and GPAI obligations already apply.
What share of pharma AI pilots reach production?
There is no clean pharma-specific figure. Deloitte State of AI in the Enterprise 2026, surveying 3,235 leaders across 24 countries, found only 25% of organisations had moved 40% or more of their pilots into production. Benchling reported 40% of biotech AI pilots reaching scaled deployment and 17% able to prove measurable value in discovery.