03Regulatory

Costing a Health Authority Query Response Before You Automate It

The industry number everyone quotes for health authority query effort comes from 12 companies surveyed in 2022, and its own arithmetic does not close. Measure your own queries first; the business case is worth more than the benchmark.

The number quoted in every health authority query automation deck comes from one place: an industry survey published by Indegene's Regulatory Affairs Digital Council on 22 August 2025. It reports roughly 200 hours per query, an average of around 70 queries per individual application, and 500 to 3,500 man-hours a year spent responding to queries on each application filing. Those are the three figures that end up on the slide.

Read the methodology and the slide gets harder to defend. The survey ran in 2022. It has 12 respondents — pharmaceutical, vaccine and biologic organisations in the US and Europe, director level and above, more than two-thirds from companies above $20bn in annual revenue. And the three headline numbers do not multiply. Seventy queries at 200 hours each is 14,000 hours per application. The survey's own annual range tops out at 3,500. Either the 200 hours covers a group of related questions rather than a single one, or the 70 queries are spread across a multi-year lifecycle while the 500–3,500 is a single year, or both. The report does not reconcile them, and neither does anyone quoting it.

This is not a reason to ignore the survey. It is a reason to treat it as what it is: a prior that tells you the order of magnitude and where to point a measuring instrument. The same dataset names the obstacle precisely. Eighty-two per cent of respondents already run a Regulatory Information Management system or similar, 80 per cent rank their own digital maturity as "Basic" on a three-point scale, and 31 per cent say a clearly defined ROI is what they need before digitising. Those three findings together describe a function that owns the tooling, knows it is not using it, and cannot write the business case. The first deliverable is therefore not a pilot. It is a number you measured yourself.

In short
  • The industry benchmark — 70 queries per application, ~200 hours per query, 500–3,500 hours per filing per year — comes from 12 companies surveyed in 2022, and 70 × 200 exceeds the survey's own annual ceiling by a factor of four to twenty-eight.
  • 30 per cent of respondents named Quality/CMC as the biggest question source. Module 3 is the reusable corpus because ICH M4Q(R1) fixes its heading tree — 72 numbered headings under 3.2 — for every product ever filed.
  • Elapsed time is set by procedure, not effort: EU applicants generally get up to three months for the day 120 list of questions and up to one month for the day 180 list of outstanding issues.
  • On 30 August 2026 the binding EU computerised-systems text is the 2011 Annex 11. Draft Annex 22's line on generative AI is consultation text, and it is scoped to critical GMP applications — not to a drafting assistant in regulatory affairs.
  • Using published RAPS compensation data, 200 hours of a US regulatory manager's time is roughly $21,000 fully loaded. That is the unit your CFO will actually price.

What the only public benchmark can and cannot support

Before the survey goes near a business case, decide what each figure is load-bearing for. Most of them are not.

Survey findingSource strengthDefensible use
~70 queries per individual applicationAttributed to a single respondent's accountSizing a measurement exercise; not a target or a denominator
~200 hours per query, coordinating, preparing and authoringAggregate across 12 organisations, unit ambiguousOrder-of-magnitude for effort per query cluster
500–3,500 hours per year per application filingAggregate range, wide by a factor of sevenThe honest headline: nobody knows this to better than a factor of seven
30% name Quality/CMC as the biggest source12 respondents, categorical questionWhere to point the pilot. Holds up because it is a ranking, not a magnitude
82% run a RIM system, 80% self-rate "Basic"Categorical, consistent with field experienceThe buying story: the platform is bought, the capability is not
31% need a defined ROI before digitisingCategoricalThe reason your first invoice is for measurement

Rankings survive a small sample far better than magnitudes do. Thirty per cent of a 12-person panel is three or four people, which is thin, but it is corroborated by where the procedural friction sits. The survey separately records FDA and EMA as the primary bodies involved, with more than 50 queries each in scope.

Why the clock, not the hours, is what gets funded

Effort in hours is a cost line. Elapsed time is a revenue line, and it is the one a commercial sponsor will fund. The two are governed by completely different things, and the second is written down in procedure.

In the EU centralised procedure, the EMA's own step-by-step description sets the shape. At day 120 the CHMP adopts a list of questions and the clock stops; the developer "generally has up to three months to answer the list of questions", with the duration agreed by the CHMP. At day 180 an updated assessment report carries a list of outstanding issues and the clock stops again; here the developer "generally has up to one month". By day 210 of active evaluation time at the latest, the CHMP adopts an opinion.

That asymmetry is the operational story. The day 120 response is a three-month programme with a project plan. The day 180 response is a four-week sprint on the questions that survived the first round, which are by construction the hardest. Automation that only shortens the first is optimising the part that was not the constraint.

The FDA side has no equivalent clock stop. Information requests are issued through the review cycle and carry no fixed statutory response deadline, but the review clock does not stop while you draft. The pressure comes from the amendment rules instead: under the PDUFA VII commitment letter, covering fiscal years 2023 to 2027, a major amendment to an original application or efficacy supplement submitted during the review cycle may extend the goal date by three months, and a major amendment to a manufacturing supplement by two months, with one extension per review cycle. A slow, incomplete response that eventually arrives as a major amendment costs more than hours. It moves the action date.

There is a proposed change worth tracking, and it is proposed, not binding. FDA published the proposed PDUFA VIII commitment letter on 14 August 2026, covering fiscal years 2028 to 2032, with a hybrid public meeting on 16 September 2026 and written comments due 16 October 2026. Among the communication commitments, FDA states an intention to include a description of the issue that triggered an information request, for IRs issued after a late-cycle meeting. If that survives negotiation and enactment, the triage step — working out what the reviewer is actually worried about — gets cheaper for a subset of IRs from FY2028. Nothing about it is in force today.

How to build the baseline in 3 weeks

The measurement is not hard. It is just nobody's job. Five fields per query, backfilled across the last 24 months, from records you already hold.

FieldWhere it livesReliability
Receipt date and submission dateRIM correspondence records, eCTD sequence metadataHard. Timestamped, auditable
Health authority and procedure stepRIM, submission planHard
Dossier location, to the ICH M4Q headingThe response document itself; often only inferableMedium. Needs a pass by a human
Named contributorsEmail and document version historyMedium
Hours expendedSelf-report or timesheetSoft. Label it as soft in the business case

Four of those five are recoverable from systems of record. Only the fifth is not, and this is where internal business cases quietly break: they present a self-reported hours figure with the same confidence as a timestamped date. Elapsed time from receipt to submission is a fact; effort in hours is an estimate made by people with an interest in the answer. Report them differently, and put the elapsed-time distribution — median, 90th percentile, and the tail of queries taking more than 60 days — in front of the sponsor first. The tail is where the money is.

Two weeks of a regulatory operations analyst plus one week of a senior reviewer to classify dossier location covers a portfolio of a few hundred queries. That is the invoice that answers the 31 per cent.

Why Module 3 is the corpus

Quality and CMC generating 30 per cent of questions is not the reason to start there. The reason is structural: Module 3 has a fixed shape. ICH M4Q(R1), current Step 4 version dated 12 September 2002, defines the Common Technical Document quality module heading by heading. Counting the numbered headings under section 3.2 in that version on 30 August 2026 gives 72 — from 3.2.S.1.1 Nomenclature through 3.2.S.4.5 Justification of Specification, 3.2.P.2 Pharmaceutical Development, 3.2.P.8 Stability and the 3.2.A appendices. Every product ever filed under the CTD uses those same headings. A question about justification of an impurity specification lands at 3.2.S.4.5 whether the molecule is a small-molecule oncology asset or a peptide.

That is what makes the corpus reusable across drug classes in a way a clinical query corpus is not. A clinical question is bound to a protocol; a Module 3 question is bound to a heading, and the heading is the same everywhere. It is also why the Quality Overall Summary, which M4Q says should normally not exceed 40 pages of text, or 80 for biotech and more complex processes, is the natural retrieval index: it follows the outline of Module 3 and contains nothing that is not already in Module 3.

Be precise about what is reusable, because this is where pilots overclaim. What transfers across products is argument structure and precedent: how you framed a mutagenic impurity control strategy, which ICH guideline you anchored to, what the assessor accepted last time. What does not transfer is a single number, batch or specification limit. A system that retrieves the shape of a previous answer and hands it to an author is doing something useful. A system that retrieves a previous answer's content and drops it in is manufacturing a data integrity finding.

The specific failure mode to design against: a superseded commitment. Your 2023 answer on a stability protocol may have been overtaken by a 2025 variation. Retrieval that ranks by semantic similarity will happily surface the older, better-written one. The corpus needs a supersession flag before it needs a better embedding model.

Which rules actually apply to a drafting assistant

Get this wrong and the project stalls in quality review for a quarter. Get it right and it is a short conversation.

On 30 August 2026 the binding EU text for computerised systems in GMP is Annex 11 (2011). The revised Annex 11 and the new Annex 22 on artificial intelligence were published for public consultation on 7 July 2025 alongside a revised Chapter 4; the European Commission consultation page records the response period as closed on 7 October 2025 and lists no successor text. Draft Annex 22's statement that generative AI and large language models should not be used in critical GMP applications is therefore consultation text, scoped to AI with direct impact on patient safety, product quality or data integrity. A tool that drafts a regulatory affairs response for a named human to author, review and sign is neither in force nor in scope. Say both halves of that sentence in the quality meeting.

FDA's framework is a different question, and mostly a scoping one. The January 2025 draft guidance on the use of artificial intelligence to support regulatory decision-making for drug and biological products sets out a seven-step credibility assessment for a model in a defined context of use, with comments closed on 7 April 2025; I could not verify a final version as of 30 August 2026, so treat the draft as the operative public reference. The scoping point matters more than the steps: the framework addresses models whose output supports a regulatory decision about safety, effectiveness or quality. A retrieval assistant surfacing your own prior correspondence produces no such output. The boundary is exactly where the FDA credibility framework stops and regulatory operations begins, and knowing which side you are on saves a credibility assessment plan you do not owe.

The EU AI Act is settled enough to state plainly. Regulation (EU) 2026/1744 of 8 July 2026, the Digital Omnibus on AI, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026. It moved standalone Annex III high-risk systems to 2 December 2027 and product-embedded Annex I systems to 2 August 2028. A HAQ drafting assistant is very unlikely to be Annex III high-risk; those categories cover employment, credit, biometrics and essential services. Live today and applicable: prohibited practices, the Article 4 AI literacy duty, general-purpose AI model obligations and Article 50 transparency. Article 4 is the one regulatory affairs functions keep missing, because it attaches to the deployer, not the vendor.

Where the output goes on to become product information rather than correspondence, the analysis changes and tightens considerably — that is the territory covered by what EMA permits when AI drafts labelling text, and a query response that proposes an SmPC wording change crosses into it.

What this means in practice

Price the unit before you price the platform. Take the RAPS 2024 Global Compensation and Scope of Practice Report, which drew 1,961 completed submissions from more than 54,000 invitations sent in March 2024, a 3.6 per cent response rate. It puts average total compensation for a US Senior Manager/Manager in the regulatory profession at $168,328, working 45 hours a week; a Specialist/Associate/Coordinator at $114,963 and 42 hours; a VP/Senior Director/Director/Associate Director at $293,251 and 47 hours. In Europe, the equivalent Senior Manager/Manager total compensation is €95,427.

The arithmetic below is mine, not RAPS's, and the assumptions are stated so they can be argued with. Take the US manager at 45 hours a week over 46 working weeks: about 2,070 hours, so roughly $81 an hour of total compensation. Add 30 per cent for employer overhead and you are near $106 an hour fully loaded. Those 200 hours are then about $21,000 per query at manager level, and a query group pulled together by a director and two SMEs costs more. Apply the same rate to the survey's annual range and one application filing consumes $53,000 to $371,000 a year in query response. Apply it to 70 queries at 200 hours and you get $1.5m, which is precisely why the survey's own arithmetic needs reconciling before anyone puts it in a deck.

Now size the prize honestly. If a drafting assistant removes 20 per cent of the authoring effort on the 30 per cent of queries that are Quality/CMC, on a 70-query application: 70 × 0.3 × 200 × 0.2 is 840 hours, or roughly $89,000 per application. That is a real number with visible assumptions, and it is defensible in a way that a vendor's percentage is not. It is also not a cycle-time saving. If the constraint is a subject matter expert with four other applications and a manufacturing investigation, faster drafting produces a first draft that waits longer. Measure the queue before you buy the drafting.

On Monday: pull two years of query correspondence out of RIM, classify each one to an M4Q heading, and plot elapsed days by heading. You are looking for the two or three headings that produce both the highest volume and the longest tail. That is a one-slide answer and it will be more specific than anything in this article, because it will be about your dossiers.

Who signs it: the head of regulatory operations owns the baseline; the head of CMC regulatory owns the corpus scoping decision and the supersession rules; quality owns the position that the system is not a critical GMP application, in writing, before build starts, citing the 2011 Annex 11 as the binding text and the July 2025 drafts as drafts. If the pipeline ever starts producing content that supports a regulatory decision rather than correspondence, re-run the scoping test that decides whether a credibility assessment is owed at all.

What goes wrong: self-reported hours arrive inflated and unfalsifiable, so anchor on elapsed days. The corpus turns out to live in individual mailboxes rather than RIM, which is a records management problem masquerading as an AI problem and needs solving first. And the pilot demonstrates faster drafting on a sample of easy queries, because the hard ones were tied up in SME review and never made it into the sample. Insist the evaluation set is drawn from the tail, not the median.

Questions people ask about this

How many health authority queries does a marketing application receive?
The only public benchmark is an Indegene Regulatory Affairs Digital Council survey of 12 pharmaceutical, vaccine and biologic organisations, fielded in 2022 and published on 22 August 2025. It reports an average of around 70 queries per individual application, with one respondent exceeding 200 queries a year. With 12 respondents, treat it as a prior for sizing a measurement exercise, not as a benchmark you can defend in a business case.
How long does it take to respond to a health authority query?
Effort and elapsed time are different numbers. The Indegene survey reports roughly 200 hours of coordination, preparation and authoring per query. Elapsed time is set by the procedure: in the EU centralised procedure the applicant generally has up to three months to answer the day 120 list of questions and up to one month for the day 180 list of outstanding issues. FDA information requests carry no fixed statutory response deadline.
Which part of the dossier generates the most health authority queries?
Quality and CMC. In the Indegene survey, 30 per cent of respondents named Quality/CMC as the functional area producing the most questions, ahead of clinical trial design at 24 per cent and clinical safety at 18 per cent. That matters for automation because Module 3 has a fixed heading structure under ICH M4Q(R1), so question shapes recur across products in a way clinical questions do not.
Does draft Annex 22 prohibit using an LLM to draft query responses?
No, on two counts. Annex 22 is a consultation draft published on 7 July 2025 with the response period closed on 7 October 2025 and no adopted successor text, so it is not binding. Its scope is AI in critical GMP applications with direct impact on patient safety, product quality or data integrity. A drafting assistant whose output is authored, reviewed and signed by a named regulatory professional is not that.
What should the first deliverable of a HAQ automation project be?
A measured baseline of your own query traffic, not a pilot. Capture receipt date, submission date, health authority, dossier location to the ICH M4Q heading, and named contributors for every query in the last 24 months. Elapsed time is measurable from records you already hold. Effort in hours is self-reported and should be labelled as such in any business case.