eCTD 4.0 in Japan, Europe and America — and Why Publishing Should Stay Deterministic
Three authorities are on three different clocks, and the gate they all put in front of you is a fixed list of machine-checked rules. That is the one place in regulatory operations where a language model has nothing to add.
Three major authorities are running three different eCTD v4.0 clocks, and only one of them has actually closed. PMDA has accepted v4.0 since April 2022 and set the end of its transition period at March 2026, with the full move to v4.0 scheduled from April 2026 — the schedule published in PMDA's own implementation briefing reads "eCTD v4 に完全移行するのは2026年4月~を予定". EMA opened optional use for new centrally authorised product marketing authorisation applications on 22 December 2025 and requires applicants to email its eCTD v4.0 team before sending anything. FDA has accepted v4.0 voluntarily since 16 September 2024 and has named no mandatory date.
That spread matters less than what sits at the end of all three pipes. Every one of these authorities checks an incoming package against a published, versioned, numbered list of rules before a human assessor sees it. The EU list for v4.0 contains 74 criteria. Japan's contains more than 250. The rules are things like maximum path length in characters and whether a UUID is well formed under ISO/IEC 11578. They are deterministic predicates over XML and a folder tree, and the result is pass or fail.
This is the single point in regulatory operations where a large language model is the wrong instrument. Not risky, not premature — wrong in the way a thermometer is the wrong tool for measuring length. A model returns a distribution over outputs; the gate returns a bit. Putting the first in front of the second adds variance to a process whose entire value is that it has none.
- Japan is the only region where v3.2.2 has closed for new applications: PMDA's transition period ran to March 2026, with full v4.0 operation from April 2026 under PSEHB/PED Notification No. 0705-1, last revised 10 March 2025.
- EU v4.0 is optional only. EMA's eSubmission pages, consulted 30 August 2026, show strongly recommended use at Q1 2027 and mandatory at Q1 2028 — a slip from the 2027 mandate foreseen in May 2025.
- FDA has published specifications but no mandatory date. The 2029 cutover quoted everywhere is an observer's projection, not an FDA commitment.
- The EU eCTD v4.0 Validation Criteria v1.1, final March 2026 and applicable from 15 July 2026, contains 74 rules: 55 pass/fail and 19 best practice. That is the whole machine gate for a v4.0 CAP application.
- The often-quoted "15% fail gateway validation" has no traceable primary source. The only sourced figure is a sub-2% FDA rejection rate given by a named FDA speaker at DIA RSIDM in 2023.
Which timeline is binding, and which is a plan
The distinction that matters when you are budgeting is between a rule that has taken effect, a date an authority has published as its intention, and a date the trade press has inferred.
| Authority | Status on 30 Aug 2026 | Key dates | Instrument |
|---|---|---|---|
| PMDA (Japan) | v4.0 required for new applications | Accepted from Apr 2022; transition to Mar 2026; full v4.0 from Apr 2026 | PSEHB/PED Notification 0705-1 (5 Jul 2017, rev. 10 Mar 2025) |
| EMA (EU) | Optional, by prior arrangement | Optional from 22 Dec 2025; recommended Q1 2027; mandatory Q1 2028 | EU eCTD v4.0 IG v1.2 + Practical Guidance v1.0 (Dec 2025) |
| FDA (US) | Voluntary, no mandate announced | Voluntary from 16 Sep 2024; Regional M1 IG v1.8 (20 Oct 2025) | Specifications issued under FD&C Act s.745A(a) |
Two of those rows have moved recently in ways that a 2025 plan will have got wrong. In May 2025 the DIA Global Forum reported EU mandatory use as "currently foreseen for 2027"; EMA's own eSubmission pages now show Q1 2027 as strongly recommended and Q1 2028 as mandatory. And the Japanese carve-out is easy to miss: PMDA's briefing states that an application submitted in v3.2.2 stays in v3.2.2 to the end of its lifecycle. There is no retroactive conversion, which means most companies will run both formats in parallel for years, not months.
Meanwhile the v3.2.2 rules underneath are still moving. The eSubmission Expert Group's implementation timeline withdrew EU validation criteria v7.1 at close of business on 30 November 2025, and since 1 December 2025 "only eCTDs compliant with EU M1 v3.1.1 and validation criteria v8.2 are accepted". A team that pinned its validator to v7.1 and stopped watching now fails at the door.
What v4.0 actually changes inside the package
The headline is that folder structure stops carrying meaning. In v3.2.2 the hierarchy encoded the dossier; in v4.0 the meaning lives in the XML message, in keywords drawn from controlled vocabularies and in UUIDs that tie a document to the contexts in which it is used.
The EU Practical Guidance v1.0, published December 2025, spells out the resulting submission unit: a first-level folder named for the EMA product number, a second-level folder that is the sequence number between 1 and 999999, a submissionunit.xml message, a sha256.txt file carrying the checksum of that message, and the module folders. Module 1 becomes "a single folder with no additional folder structure", with country, dosage and strength expressed as keywords rather than directories.
PMDA put the operational consequence more bluntly in its briefing slides: compared with v3.2.2, grasping the structure by human eye is difficult, and tool support is required. The visual QC pass that experienced publishers have relied on for fifteen years does not survive the format change. What replaces it is the validator, which makes the validator's correctness and version control a first-order compliance concern rather than an IT detail.
How many machine checks stand between you and the assessor
Fewer than most people assume, and they are all published.
Downloading the EU eCTD v4.0 Validation Criteria v1.1 — the final version dated March 2026, applicable from 15 July 2026 — and counting the rows on 30 August 2026 gives 74 numbered criteria, of which 55 carry severity P/F and 19 carry BP. Criterion eCTD4-EU-001 requires that the UUID for a regulatory activity be identical across every submission unit referring to it. Criterion eCTD4-EU-072 requires every UUID to be well formed under ISO/IEC 11578:1996 and ITU-T Rec X.667. These are not judgement calls.
Japan's list is longer and more physical. The eCTD v4.0 JP Check Items List v1.6.0.0, English provisional translation dated March 2025, runs to 70 pages; extracting its tables on 30 August 2026 yields 252 live check rows, with identifiers running up to JP-eCTD4-362 because retired items keep their numbers. Item JP-eCTD4-018 caps the path from the first-level folder at 180 characters. JP-eCTD4-020 caps a dossier folder name at 64. JP-eCTD4-023 caps a dataset filename at 32. JP-eCTD4-024 requires each file name to contain exactly one extension. The v1.6.0.0 revision of 26 March 2025 did nothing more exotic than add the condition that a controlled-vocabulary code must be active, not merely valid, across twelve check items.
FDA's equivalents are its Specifications for eCTD Validation Criteria and, for study data, the Technical Rejection Criteria. Validation codes 1734, 1735, 1736, 1737 and 1789 took effect on 15 September 2021 after a six-month warning period that began on 15 March 2021, and code 1789 — a file submitted in a study section without an accompanying study tagging file — carries high severity, meaning the submission is not received.
That last point is a regional trap. Japan and the United States do not treat a validation failure the same way. PMDA's briefing states plainly that validation results do not affect receipt of the application; at FDA, a high-severity failure means the submission never lands. A single global "first-pass rate" KPI that averages across both is measuring two different things.
Do 15 per cent of submissions really fail gateway validation?
This number circulates widely in vendor content and in the AI-generated regulatory blogs that now dominate the search results for eCTD validation. I could not trace it to a primary source, and I would not put it in a business case.
What does survive checking is smaller and better attributed. Ennov's published analysis of FDA eCTD rejection reasons attributes a rejection rate of "less than 2%" to a 2023 DIA RSIDM presentation by Ethan Chen, then FDA's Director of Data Management Services and Solutions, and reproduces his chart of the most common causes. Ennov's own conclusion is the useful part: the listed errors "should be identified by any competent validation tool", which points at either a validator defect or, more often, a process that changed the package after the last validation run.
So the honest framing is not a failure rate. It is a failure mode: submissions fail on things a machine already knows how to catch, at organisations that either ran the check too early or ran it against the wrong criteria version. Both are configuration problems. Neither is an intelligence problem.
Why a language model is the wrong tool at the gate
Take the four properties of the validation step together. The rules are published in advance. They are expressed as predicates over structured data — a character count, a regular expression, a UUID format, a cross-reference. The outcome is binary. And the authority runs the same check you can run, using a specification you can download.
A system with those four properties has a reference implementation available to both parties. The correct architecture is a rules engine that encodes the published criteria, versioned to the criteria version, with a test suite of known-good and known-bad packages. When PMDA adds the "code must be active" condition to twelve check items, you change twelve rules and re-run the suite. The change is reviewable in a diff, which is what an inspector will ask to see.
A generative model in that position has no advantage to trade against its costs. It cannot be more correct than the published rule, because the published rule is the ground truth. It introduces run-to-run variance into a step whose whole purpose is to have none. It makes the qualification exercise dramatically harder, because you are now demonstrating consistent behaviour of a stochastic component in front of a gate that a deterministic component satisfies by construction. And in the failure case it is worse than useless: a model that plausibly asserts a package is valid, when the validator disagrees, has cost you the submission window.
The line is not "no AI in regulatory operations". It is that the line sits upstream of the validator. Drafting Module 2 summaries, reconciling a submission plan against a dossier index, extracting commitments from a health-authority letter, triaging which of 300 documents in a legacy dossier need remediation before conversion — these are open-ended language tasks with a human reviewer downstream and no binary machine oracle. That is where a model earns its keep. Publishing and validation is where it does not.
Where the regulatory line sits for the AI you do use
If you put a model anywhere in the submission chain, know which instrument binds you today.
In the EU, the binding text for computerised systems in GMP on 30 August 2026 is still Annex 11 of 2011. The draft revised Annex 11 and the new draft Annex 22 on artificial intelligence were published for public consultation on 7 July 2025, the consultation closed on 7 October 2025, and neither is law. Draft Annex 22's position that generative AI and large language models are not acceptable in critical applications is a strong signal of inspector thinking and a sensible planning assumption, but it is not currently an enforceable requirement, and writing it into an SOP as though it were is the kind of error that ends a credibility conversation with a QA director. For clinical systems, the applicable EMA text is the guideline on computerised systems and electronic data in clinical trials, effective 10 September 2023.
On the EU AI Act, the position changed this summer and much published commentary is stale. The AI Omnibus was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026. It defers the high-risk obligations for Annex III systems to 2 December 2027 and for AI embedded in products under Annex I to 2 August 2028. This is enacted law, not a proposal. The prohibitions on unacceptable practices and the obligations on general-purpose AI models were not deferred and already apply.
A deterministic eCTD validator is not, on any reading, a high-risk AI system — it is not an AI system at all. That is a point in its favour that rarely makes it into the architecture discussion.
What this means in practice
Pin your validator to a criteria version and treat that pin as a change-controlled configuration item, not a tool setting. The EU withdrew v7.1 overnight on 30 November 2025 and Japan re-issued its check items in March 2025; both changes are trivial to absorb if you know which version you are on and painful if you do not.
Run the final validation after the last change to the package, immediately before transmission, and record the result with the sequence. This is the fix for the single most common documented failure mode, and it costs nothing but a step in the SOP.
Split the two programmes in your budget. Outsourced eCTD publishing runs roughly $5,000 to $20,000 per submission and $10,000 to $50,000 for a full application, with building the capability in-house estimated at $300,000 to $500,000 in the first year over 9 to 18 months — vendor-side figures published by IntuitionLabs and by Extedo, so treat them as order-of-magnitude rather than benchmarks. A generative-AI authoring pilot is a separate line with a separate owner and a separate validation argument. Merging them into one "AI for regulatory" business case is how a publishing programme ends up carrying a model qualification it never needed.
Decide now who signs the v4.0 dual-running plan. Most companies will file in v4.0 in Japan, v3.2.2 in Europe and either in the United States for several years, on three criteria versions that change independently. That is a configuration-management problem owned by regulatory operations, and it is the work that actually determines whether your submissions land. The interesting AI question is upstream, in the documents themselves — and it should stay there.
Questions people ask about this
- When did eCTD v4.0 become mandatory in Japan?
- PMDA began accepting eCTD v4.0 in April 2022 and set a transition period running to March 2026, with the full move to v4.0 scheduled from April 2026. The underlying instrument is PSEHB/PED Notification No. 0705-1 of 5 July 2017, revised most recently on 10 March 2025. Applications already filed in v3.2.2 continue in v3.2.2 through to the end of their lifecycle.
- Is eCTD v4.0 mandatory in the EU yet?
- No. Optional use for new centrally authorised product marketing authorisation applications went live on 22 December 2025, and applicants must contact the EMA eCTD v4.0 team before submitting. EMA's eSubmission pages, consulted on 30 August 2026, put strongly recommended use at Q1 2027 and mandatory use at Q1 2028. eCTD v3.2.2 remains accepted throughout.
- Has FDA set a date for mandatory eCTD v4.0?
- Not as of 30 August 2026. CDER and CBER have accepted v4.0 voluntarily since 16 September 2024, and FDA has continued to publish v4.0 specifications, including Regional Module 1 Implementation Guide v1.8 dated 20 October 2025. The 2028 to 2029 cutover quoted across industry commentary is a projection by observers, not a date FDA has committed to.
- What proportion of eCTD submissions fail validation?
- There is no reliable public figure. The only number traceable to a named FDA speaker is a rejection rate below 2 per cent, given by Ethan Chen of FDA at DIA RSIDM in 2023. The widely repeated claim that roughly 15 per cent of submissions fail gateway validation has no primary source that survives checking, and the regions do not even define failure the same way.
- Should AI be used for eCTD publishing and validation?
- Validation is a fixed list of machine-checkable rules with a binary outcome, so it is a rules-engine problem and a generative model adds variance without adding information. The defensible place for a language model is upstream authoring and pre-submission triage, where a qualified human reviews the output before it reaches a validator.