03Regulatory

Ninety-Seven Percent Faster, Seventy Percent Complete

A generative model drafted three eCTD nonclinical summaries from 18,870 pages in 3.7 hours. The same paper reports that it assigned NOAEL values to in-vitro studies and omitted a GLP-required analysis from every summary it wrote. Both numbers are load-bearing.

Two large-molecule INDs, already cleared by FDA, were run back through a generative drafting tool to see how much of the original writing effort it could reproduce. For the first, AutoIND turned 61 nonclinical study reports totalling 18,870 pages into complete first drafts of eCTD Modules 2.6.2, 2.6.4 and 2.6.6 in 3.7 hours. For the second, 58 reports and 11,425 pages took 2.6 hours. Against an estimated 100 hours of manual drafting, that is a 96.3% and a 97.4% reduction, and it moves throughput from 0.2 to 12.1 pages per hour.

The quality half of the same paper is where the interesting reading is. A single experienced regulatory writing assessor scored the output across seven dimensions and reached 69.58% and 77.85% overall. Correctness held above 89% in every module. Prominence and emphasis in the toxicology summary scored 6.7% for the first IND. Completeness in toxicology reached 32%. The tool omitted the GLP-required dose-formulation analysis from 100% of the summaries it produced, left 35% of summaries missing essential study design elements, and assigned NOAEL values to in-vitro studies, which have no NOAEL to assign.

Read those two halves together and the finding is narrower than the headline and more useful. The model performed compilation: it read the source reports and reproduced their facts with high fidelity. What it could not do was decide which facts mattered. Compilation is roughly 70% of the work by volume and close to none of it by consequence, because an ICH-compliant nonclinical written summary is a document whose whole function is selection under a page ceiling.

In short
  • AutoIND drafted three nonclinical written summaries from 18,870 pages in 3.7 hours, versus an estimated 100 hours manual. The comparator was never measured.
  • Overall quality was 69.58% and 77.85%; correctness stayed above 89%, but prominence and emphasis in toxicology fell to 6.7%.
  • The paper reports a 100% omission rate for the GLP-required dose-formulation analysis and NOAEL values assigned to in-vitro studies, while its abstract states no critical regulatory errors were detected. Its own protocol had named both as examples of a critical error.
  • Output ran 3 to 5 times longer than expert-written text, against an ICH M4S recommendation that all three summaries together stay inside 100 to 150 pages.
  • On 30 August 2026 this is an arXiv preprint, not peer-reviewed, funded jointly by Takeda and Weave, with a single assessor and two INDs from one company and one modality.

What the study actually measured

The work was done by authors at Weave Platform and Takeda Pharmaceuticals and posted as arXiv:2509.09738 on 10 September 2025. Checked against its arXiv listing on 30 August 2026, it remains a single-version preprint with no journal reference: it has not been peer reviewed. Funding came jointly from Takeda and Weave, and the Weave authors declare that they are employees and shareholders of Weave. The one countervailing detail worth noting is that the expert IND analysis was performed by an author declared as a consultant and not a shareholder of either company.

The configuration matters for anyone reading the result forward. AutoIND V2.3 ran gpt-4-turbo, release 2024-10, inside an AWS VPC, with source PDFs extracted by AWS Textract, classified into IND sections and inserted into prompts tuned to Takeda's style guide. Five reports from each IND were used to calibrate those prompts and then excluded from scoring, which is the right call methodologically and also means the reported performance is post-tuning, not out-of-the-box. The drafts were never submitted to FDA.

IND-1 was a large molecule submitted in June 2021, IND-2 a large molecule submitted in January 2024. Both were manually written, submitted and cleared. So the source material is real, the task is real and the timing is measured with a stopwatch — drafting times were logged in Toggl, including PDF upload and content extraction. That last inclusion is more honest than most vendor benchmarks, which quietly exclude ingestion.

Where the 97% came from, and what the 100 hours were

The 3.7 hours is measured. The 100 hours is not. The paper is explicit: manual drafting times "were estimated based on the experience of regulatory writers" with at least six years' experience at Takeda, and are used "only as industry-standard benchmarks to contextualize efficiency". There was no parallel human arm, no time-and-motion study of the original 2021 and 2024 drafting effort, and no reconciliation against the actual timesheets for those submissions.

That does not make the number wrong. It makes the denominator a professional judgement, and a 97% reduction computed against a judgement is a different object from a 97% reduction computed against a control. The reported effect size of d = 6.8 with p < 0.001 is arithmetic on two measured values against an estimated constant, not evidence of a between-arms difference. This is the recurring problem with published acceleration figures across the sector, and it is worth reading alongside how disclosed AI cycle-time numbers compare once you check what each one is measuring before any of them go into a business case.

The paper's own limitations section concedes the deeper point: real documents go through multiple revision cycles, and "a more comprehensive assessment would compare final polished documents from both approaches". The 97% is a first-draft metric. Nobody has published the number that a regulatory operations director actually needs, which is time from kick-off to a summary the nonclinical lead will sign.

The 7 quality dimensions, scored

Each dimension comprised multiple sub-criteria scored 0 to 3, summed and normalised to the category maximum. The percentages are not relative to human performance; they are relative to submission-ready expectation.

DimensionIND-1IND-2Reported defect pattern
Correctness>89% (all modules)>89%Dose-response claimed where absent; latency confused with duration
Completeness, 2.6.6 toxicology32%59%Missing N/group, vehicle composition, dose regimens
Completeness, 2.6.4 PK61%76%Absent toxicokinetics, control results, IC₅₀/EC₅₀ values
Conciseness, 2.6.4 PK30%66%3 to 5 times word-count inflation
Prominence/emphasis, 2.6.66.7%37%Method-heavy, result-thin in over 60% of mentions
Overall69.58%77.85%

Clarity and redundancy were described as relatively strong and consistent, with no per-module figures given in the text. The clarity breakdown that is given is instructive: of 156 clarity comments, 32% concerned structure and organisation, including results placed before methods, and 28% concerned AI-characteristic language, with the phrase "The study aimed to evaluate" counted 12 times.

Two structural weaknesses limit how far these percentages travel. The scoring was done by a single assessor, which the authors acknowledge "introduces potential bias", and they note that the most subjective dimensions are the two that carry the argument — clarity and emphasis. A 6.7% emphasis score is a strong signal, but it is one person's rubric applied once, and the difference between 6.7% and 37% across the two INDs is larger than any difference in the tool. The second weakness is generalisability: two INDs, both large molecules, one therapeutic area, one company, one house style guide. Nothing here tells you how the same tool performs on a small molecule with a busy genotoxicity package, and the authors say so.

One arithmetic snag is worth flagging because it will be repeated. The paper states that "10-25% of generated content required revision or refinement", immediately after reporting 69.58% and 77.85%. The complements of those scores are 30.4% and 22.2%. The stated revision range does not follow from the stated scores, and the higher figure is the one that should go into a resourcing model.

Why emphasis collapsed to 6.7%

ICH M4S(R2) (CPMP/ICH/2887/99, Step 5) says the primary purpose of the nonclinical written summaries is "a comprehensive factual synopsis of the nonclinical data", with interpretation, clinical relevance and labelling implications belonging in the 2.4 Nonclinical Overview. On a naive reading, that sounds like exactly the compilation task an LLM is good at.

The guideline's operative clauses say otherwise. It recommends that the three written summaries together "in general not exceed 100-150 pages". The toxicology brief summary should run to "a few pages (generally not more than 6)". Section 2.6.6.3 asks that repeat-dose studies be summarised "giving brief details of the methodology and highlighting important findings", and that non-pivotal studies "be summarized in less detail" than the definitive GLP studies specified by ICH M3. Every one of those is an instruction about proportion. Under a page ceiling, deciding what to compress is the writing.

That is precisely the axis on which the model failed, and it failed systematically rather than randomly. Emphasis scored low in every module and both INDs, worst in toxicology, where the source reports are longest and most method-dense. The model over-described procedures — the paper names the pattern the "methods novella", including buffer recipes and equipment specifications — and under-weighted results. Sixty to seventy percent of cases needed structural repositioning; 70 to 80% needed content reduction. A tool that inflates text 3 to 5 times while being asked to prioritise under a page budget is not making a small error at the margin. It is optimising the wrong objective.

2 findings that sit awkwardly against "no critical regulatory errors"

The paper predefines a critical regulatory error as "any misrepresentation or omission likely to alter regulatory interpretation of safety, efficacy, or compliance", and offers exactly two worked examples: "incorrect NOAEL attribution" and "omission of mandatory good laboratory practice (GLP) dose-formulation analysis".

Both examples then appear in its own results. Under correctness, the paper reports that the model "inappropriately assigned NOAEL (No Observed Adverse Effect Level) values to in-vitro studies". Under completeness, it reports a "100% omission rate for GLP-required dose formulation analysis", restated in the discussion as a "100% failure rate for specific elements like dose formulation analysis in GLP studies". The abstract nevertheless states that no critical regulatory errors were detected.

These can be reconciled — the assessor may have judged that in a first draft, before human refinement, neither defect could yet mislead a reviewer, and the "e.g." examples may have been illustrative rather than determinative. But the reconciliation is not stated in the paper, and the abstract sentence is the one that will be pasted into steering-committee decks. Do not carry it forward without the caveat.

The dose-formulation omission is the more serious of the two, because it is not a stylistic gap. Under 21 CFR 58.113, a binding US regulation, any test or control article mixed with a carrier must be analysed for uniformity of the mixture, periodically for concentration and, per the conditions of the study, for stability. Those analyses are what make a GLP toxicology dose a known quantity. A 2.6.6 summary that never mentions them describes exposures a reviewer cannot verify. The NOAEL error is different in kind: it is a fabricated regulatory conclusion attached to a study design incapable of supporting one, which is the same failure mode that makes non-deterministic LLM output hard to validate under GxP — plausible form, absent referent.

Which rules apply to a tool that drafts a submission

Very little applies directly, and getting this right is the difference between a defensible position and an overclaim in front of a QA director.

InstrumentStatus on 30 Aug 2026Relevance to AI drafting of Module 2.6
ICH M4S(R2)Adopted guidelineSets the content and proportion requirements the draft must meet
21 CFR Part 58 (GLP)Binding US regulation58.113 is the source of the dose-formulation requirement omitted
FDA AI guidance, 6 Jan 2025Draft, comments closed 7 Apr 2025, not finalisedExcludes operational-efficiency uses; a drafting assistant is generally out of scope
EMA–FDA guiding principles, Jan 2026Non-binding principlesPrinciple 8 requires assessing the complete system including human–AI interaction
EMA AI reflection paperFinal 9 Sep 2024, non-bindingRisk scales with the output's influence on regulatory decisions
EU GMP Annex 11 (2011)Binding operative guidanceGMP scope; an IND authoring tool is not in it
Draft Annex 22 (AI in GMP)Consultation draft, published 7 Jul 2025, closed 7 Oct 2025GMP scope only. Do not apply it to regulatory writing

The FDA scope point deserves care because it cuts both ways. The draft guidance issued 6 January 2025, whose comment period closed on 7 April 2025 and which had not been finalised as of 30 August 2026, applies to AI models producing information or data to support regulatory decision-making on safety, effectiveness or quality, and expressly does not apply to drug discovery or to uses aimed at operational efficiencies such as internal workflows. A tool that drafts prose a human then verifies and signs sits outside the seven-step credibility framework. The document it drafts sits squarely inside the submission.

The genuinely binding EU instrument here is not a GMP annex. Under Regulation (EU) 2026/1744, the Digital Omnibus on AI, published in the Official Journal on 24 July 2026 and in force from 27 July 2026, the Annex III high-risk obligations moved to 2 December 2027 and Annex I embedded systems to 2 August 2028. Regulatory authoring is not an Annex III use in any event. What is already live and does bite is the Article 4 AI literacy duty, applicable since 2 February 2025: the writers refining these drafts need documented competence in what the tool does and where it fails.

What this means in practice

The study's most transferable output is not the 97%. It is the defect taxonomy, because the failures are systematic and therefore checkable. Build the review protocol around them rather than around a general instruction to "check the AI's work".

Four checks earn their place immediately. First, a mandatory-field gate: for every GLP study summarised, confirm the presence of dose-formulation analysis, N per group, vehicle composition and dose regimen before the draft leaves the tool. A 100% omission rate is not a review finding, it is a template defect, and it should be caught by a rule, not by a reader. Second, a study-type-to-conclusion validator: NOAEL, LOAEL and adverse-effect statements must not attach to in-vitro study records. Third, a length budget enforced against ICH M4S — set a hard word ceiling per section, given a documented 3 to 5 times inflation. Fourth, an emphasis pass owned by a named human: results before methods, specific doses in place of "at higher doses", control data present wherever a dose-response claim is made.

On sizing, use the complement of the quality score, not the paper's stated range: plan for 22% to 30% of content requiring rework, with the toxicology summary consuming most of it. On evidence, note what Principle 8 of the January 2026 EMA–FDA guiding principles asks for — a risk-based performance assessment of the complete system, including human–AI interactions. A benchmark of the model alone does not satisfy that, and neither does a 97% figure that stops at the first draft. If you are going to build a case internally, measure your own end-to-end cycle time on your own documents, the way the published cycle-time claims need unpicking before they can be compared.

Someone has to sign. There is still no regulator guideline anywhere that speaks specifically to AI-authored regulatory documents, which means accountability lands exactly where it did before: on the qualified person who puts their name on Module 2.6 and, further up, on the nonclinical lead who owns the safety narrative. That is an argument for naming the reviewer of record per section in the SOP before the tool goes live, not after the first draft lands. It is also an argument against the seductive framing in the paper's discussion, which claims AI handles "the 70% cognitive lift" of regulatory writing. That figure carries no citation anywhere in the text, and its proximity to the 69.58% quality score is a coincidence. Compilation is most of the keystrokes and very little of the liability.

Finally, put the model generation in the record. This benchmark ran gpt-4-turbo from October 2024. Two years on, the compilation half will be better and the omission half may not be, because omitting a required field is a specification failure rather than a capability one. The right question to a vendor is not what the model scores. It is which fields the system refuses to emit a draft without.

Questions people ask about this

How fast did AutoIND draft the IND nonclinical summaries?
In the Weave–Takeda preprint, AutoIND produced first drafts of eCTD Modules 2.6.2, 2.6.4 and 2.6.6 in 3.7 hours from 61 source reports totalling 18,870 pages, and in 2.6 hours from 58 reports totalling 11,425 pages. The comparator, roughly 100 hours, was an estimate from Takeda regulatory writers rather than a measured control arm.
What quality score did the AI-drafted IND summaries achieve?
A single experienced assessor scored the drafts against seven dimensions on a 0 to 3 scale, normalised to a percentage. Overall scores were 69.58% for the first IND and 77.85% for the second. Correctness exceeded 89% throughout, but prominence and emphasis in the toxicology summary scored 6.7% for the first IND and 37% for the second.
Did the AI hallucinate anything in the regulatory summaries?
Yes. The paper reports that the model assigned NOAEL values to in-vitro studies, described adverse effects where none were reported, claimed dose-response relationships that did not exist, and produced numerical errors such as an 11-hour onset in place of 7 hours. The authors classify these as correctness patterns rather than critical regulatory errors.
Is AI-drafted regulatory writing covered by the FDA AI guidance?
Not directly. FDA's draft guidance of 6 January 2025, still a draft on 30 August 2026, applies to AI models producing information to support regulatory decision-making, and excludes uses aimed at operational efficiencies such as internal workflows. A drafting assistant generally falls outside it. The submission the assistant helps write does not.
Has the AutoIND study been peer reviewed?
Not as of 30 August 2026. It is arXiv preprint 2509.09738, posted 10 September 2025, with a single version and no journal reference on its arXiv listing. It was jointly funded by Takeda and Weave, and the Weave authors are employees and shareholders of Weave.