05Commercial

What Agentic Promotional Review Automates, and Where the Human Still Signs

One number on one vendor page measures the approved corpus rather than the review process, and it says the expensive defect is sitting in the claims and reference library rather than in the reviewers.

The most interesting number in agentic promotional review is not the one about speed. On the Veeva Falcon MLR product page, fetched on 30 August 2026, a headline claim of "98% accuracy on critical issues" is supported by a single line underneath: the system "catches an average of 1.3 additional compliance gaps per already-approved page". Every other figure on that page measures the review process — 80% faster time to market, minus 22% review cycles, plus 62% first-time-right, minus 70% pre-MLR rejections. That one measures the corpus.

Read it literally and it is a statement about your library, not about your reviewers. Approved pages are pages that a named signatory examined in final form and certified. If an agent finds an average of 1.3 things wrong with each of them, then either the approved corpus carries roughly one and a third defects per page that survived human medical, legal and regulatory review, or the agent is counting things that a signatory looked at and consciously accepted. Both readings point at the same object: the claims and reference library, and the record of what was decided about it.

That matters more than which agent you buy, because every agent reads the same corpus. An MLR agent is a function of two inputs: the asset, and the approved source material it is checked against. The second input is the one you own, the one no vendor can sell you, and the one that determines the ceiling on any agent's performance. The published evidence on how language models behave against structured versus narrative sources says the gap between those two states is roughly forty accuracy points. That is the whole argument of this piece.

In short
  • The 1.3 additional compliance gaps per already-approved page figure on the Falcon MLR page carries no published denominator, adjudication method or definition of "gap". It is offered as evidence for a 98% accuracy claim, but the two measure different things.
  • In the MEDAL benchmark (Patterns, 30 March 2026), GPT-4o-mini scored 94.0% against structured guideline sources and 56.3% against narrative ones. Agent accuracy tracks the structure of your library, not the brand of the model.
  • Article 92(2) and 92(3) of Directive 2001/83/EC already require promotional documentation to be accurate, up-to-date, verifiable and sufficiently complete, with quotations and tables faithfully reproduced and precise sources indicated. That is a specification for a machine-readable claims library, written in 2001.
  • In the US the adequate provision option in 21 CFR 202.1 is still available on 30 August 2026. The rescission is RIN 0910-AJ14, with a proposed rule expected December 2026 — not law, not yet even proposed.
  • The 2024 ABPI Code contains zero occurrences of "artificial intelligence", "machine learning" or the standalone token "AI". Clause 8.1 still requires one named doctor or pharmacist to certify, and Clause 8.6 still requires the certificate to be kept for three years after final use.

What the 1.3 figure does and does not establish

Take the claim at face value first. The Falcon MLR page, which I fetched on 30 August 2026, states the figure without a denominator, without a description of how a "compliance gap" was defined, without saying who adjudicated the findings, and without saying whether the already-approved pages were sampled at random or selected. There is no methodology note or source citation anywhere on the page.

That absence is the interesting part, because the two numbers it pairs are not commensurable. A 98% accuracy figure is a rate against a labelled truth set: you need a set of pages where the correct answer is known, and you count agreement. A count of 1.3 additional findings per page is a raw yield: you count what the machine surfaced beyond what the humans surfaced. Yield is not precision. If nobody adjudicated the 1.3, some unknown proportion of it is false positives, and false positives in promotional review are expensive in a specific way — each one costs a signatory's time, and signatories are the scarce resource the whole exercise is meant to protect.

Now take the pessimistic reading seriously, because it is the one with regulatory consequences. In the United States, promotional pieces are not private documents. Under 21 CFR 314.81(b)(3)(i), an applicant must submit specimens of mailing pieces and any other labelling or advertising devised for promotion of the drug product at the time of initial dissemination, accompanied by a completed Form FDA 2253. Failure to do so renders the product misbranded under section 502(n) of the Federal Food, Drug, and Cosmetic Act. The already-approved pages an agent is finding gaps in are, for the most part, pages that were filed.

So an agent run over a live approved corpus is not a productivity exercise. It is a retrospective compliance scan that generates knowledge, and knowledge creates obligations. The first governance question of any agentic MLR pilot is therefore not "how much faster is it" but "what do we do with the findings on material that is already in the field". If you cannot answer that in writing before the pilot starts, do not start the pilot.

Which checks are mechanically decidable, and which are not

The useful way to scope an MLR agent is to sort the checks by whether a correct answer exists independently of a reviewer's judgement. This is the distinction that determines what can be automated to a defensible standard and what can only be assisted.

CheckDecidable against a source?What settles it
Claim consistent with the SmPC or approved labellingYesArticle 87(2), Directive 2001/83/EC: all parts of the advertising must comply with the particulars listed in the SmPC
Reference actually supports the sentence attached to itYes, at locator levelABPI Clause 14.2: clear references must be given for published studies
Table or figure faithfully reproduced from sourceYesArticle 92(3): faithfully reproduced, precise sources indicated
Documentation carries a drawn-up or last-revised dateYesArticle 92(1) requires the date to be stated
Dual modality of the major statement in TV advertsYes21 CFR 202.1(e)(1)(ii)(C), in force since 20 May 2024
Audio "at least as understandable" as the rest of the advertNo21 CFR 202.1(e)(1)(ii)(B) — a comparative judgement
Net impression of the piece as a wholeNoThe basis of most 2026 OPDP untitled letters

The bottom two rows are where the human keeps signing, and the gap between them and the rows above is widening rather than narrowing. FDA's Office of Prescription Drug Promotion has spent 2026 issuing letters on impressions rather than on statements. Covington's April 2026 update reports nine untitled letters and one warning letter from OPDP in the first quarter of 2026 alone. By mid-July, the FDA Law Blog counted the 20th and 21st untitled letters of 2026, including one to Viatris over TOBI PODHALER material implying the device could be used "in the car" when the approved instructions call for adequate lighting and stable conditions, and one to Sanofi over emails describing prevention of "RSV disease" where the approval covers RSV lower respiratory tract disease.

Sidley's analysis of a March 2026 untitled letter on an injectable incretin product shows how far the impression standard now reaches: FDA objected to a bright orange shirt contrasted against a dull grey one as implying superiority, and to comedic tone as suggesting competing products were unworthy of substantive discussion. No agent adjudicates that. An agent can flag that a comparative implication may exist and route the asset to a signatory faster, which is worth something, but the decision is a person's and stays a person's.

Why the defect sits in the library

The best public evidence on why an agent's performance depends on the corpus rather than the model comes from outside promotional review entirely. Wang and Chen at Johns Hopkins built MEDAL, a benchmark of roughly 21,000 question-answer pairs drawn from three evidence streams — 8,530 from Cochrane systematic reviews, 2,580 from structured American Heart Association guidelines and 10,500 from narrative clinical guideline documents — and published the results in Patterns on 30 March 2026.

The spread by source type is the finding. GPT-4o-mini reached 94.0% accuracy against the structured AHA guidelines, with precision of 1.00 and recall of 0.94. The same model dropped to 60.3% against systematic reviews and 56.3% against narrative guidelines. Claude 4.5 Sonnet reached 97.0% on the structured tasks and DeepSeek-v3 91.9%. Every model tested fell substantially when the source stopped being structured.

The variable that moved accuracy by roughly forty points was the shape of the source material, not the model. That is the number to carry into a vendor conversation, because a typical pharmaceutical claims library is not structured in the MEDAL sense. It is a set of approved assets with a reference pack: PDFs of published papers, often with a highlighted passage, sometimes with a page number, frequently with the annotation living in a separate file from the claim it supports. Ask what an agent sees when it checks a claim against that, and the answer is a narrative source. You have bought a system whose ceiling was set by your document management conventions.

Size compounds the problem. The corpus an agent has to be right about is not the assets in flight this quarter but everything still certified and in use, most of which is never deployed — the benchmarks on approved content that never reaches the field put the unused share near four fifths. Refactoring a library that large is a scoping decision before it is an engineering one.

A structured claims library is a different object. Each claim is an atomic record with its own identifier. Each record carries the exact locator in the source — page, table, figure, line — rather than a document-level citation. Each carries the evidence type, the markets in which the claim is approved, the date it was drawn up or last revised, and an expiry. That last field is not an optimisation. Article 92(1) of Directive 2001/83/EC requires promotional documentation transmitted to persons qualified to prescribe or supply to state the date on which it was drawn up or last revised, and Article 92(2) requires the information in it to be accurate, up-to-date, verifiable and sufficiently complete to enable the recipient to form their own opinion of the therapeutic value. Article 92(3) requires quotations, tables and other illustrative matter taken from medical journals to be faithfully reproduced with the precise sources indicated.

Those three paragraphs, adopted in 2001, describe a claims library with locator-level provenance and revision control. Most companies satisfy them procedurally, through a reviewer who knows where things are, rather than structurally, through a data model. Procedural satisfaction is invisible to an agent.

The practical test of whether you have the structural version is already in the UK code. Clause 18.2 of the 2024 ABPI Code requires substantiation for any information, claim or comparison to be provided as soon as possible, and certainly within ten working days, at the request of a health professional or other relevant decision maker. Clause 14.3 imposes the same ten-working-day limit on data on file. If producing substantiation for an arbitrary claim in your live promotional estate takes a person half a day of searching, your library is narrative. If it takes a query, it is structured, and an agent will perform on it the way MEDAL's models performed on the AHA guidelines.

What binds you on 30 August 2026, and what does not

Promotional review attracts more confident misstatement about regulatory status than almost any adjacent area, partly because three separate regimes are moving at once. Here is the position as of today.

InstrumentStatus on 30 August 2026Effect on promotional review
21 CFR 202.1 and 314.81(b)(3)(i)BindingAdvert content standards; Form FDA 2253 at initial dissemination
"Clear, conspicuous and neutral" final ruleBinding; FR 21 Nov 2023, effective 20 May 2024, compliance 20 Nov 2024Five standards at 21 CFR 202.1(e)(1)(ii)(A)–(E) for TV and radio
Rescission of adequate provisionNot proposed; RIN 0910-AJ14, NPRM expected Dec 2026Nothing yet; would require full brief summary in broadcast
Presidential memorandum, 9 Sept 2025Executive direction, not a ruleDrove the enforcement wave; creates no new obligation itself
Directive 2001/83/EC, Articles 87 and 92Binding as transposedSmPC consistency; accurate, verifiable, dated documentation
2024 ABPI CodeBinding on members via PMCPACertification, substantiation, retention
EU AI Act Article 50Applies since 2 Aug 2026Marking of AI-generated synthetic content
EU AI Act Annex III high-riskDeferred to 2 Dec 2027 by Reg. (EU) 2026/1744Promotional review is not an Annex III use case

Three of those rows deserve expansion, because they are the ones most often stated wrongly.

The adequate provision rescission is not law and is not yet even a proposed rule. On 9 September 2025, following a presidential memorandum of the same date directing HHS to increase the risk information carried in prescription drug advertising and directing FDA to enforce the advertising provisions of the Federal Food, Drug, and Cosmetic Act, FDA and HHS announced a three-part programme: rulemaking to rescind the adequate provision option, an enforcement wave of roughly 100 letters, and expanded oversight of social media promotion including influencer partnerships, algorithm-driven targeting and AI-generated health content. The rulemaking is listed on the 2026 Unified Agenda as RIN 0910-AJ14 with a notice of proposed rulemaking expected in December 2026, designated economically significant. Anyone telling you today that broadcast adverts must carry the full brief summary is describing a possible 2027 or 2028 state, not the current one. What did change, and is binding, is the clear, conspicuous and neutral final rule published in the Federal Register on 21 November 2023, effective 20 May 2024, with a compliance date of 20 November 2024.

The AI Act's high-risk deferral is settled law, and it does not cover you anyway. Regulation (EU) 2026/1744, the Digital Omnibus on AI, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026. It moved the compliance date for standalone Annex III high-risk systems from 2 August 2026 to 2 December 2027, and for AI embedded in products covered by EU product-safety law to 2 August 2028. Promotional review is not an Annex III use case, so the deferral is mostly beside the point for MLR. What does apply is Article 4 on AI literacy, live since 2 February 2025, and the Article 50 transparency obligations, which applied from 2 August 2026 and were deliberately left out of the deferral.

Annex 11 is not your instrument. Promotional review is not a GMP activity, so the 2011 Annex 11 to the EU GMP guide does not bind it, and neither would draft Annex 22 or the draft Annex 11 revision if they were adopted — both were published for consultation on 7 July 2025, that consultation closed on 7 October 2025, and as of 30 August 2026 neither is law. This gets confused constantly because the same vendors sell into both estates. The predicate rules that make your promotional records regulated are the advertising rules and, in the US, the submission requirement at 314.81(b)(3)(i), which is what pulls electronically maintained promotional records into the scope of 21 CFR Part 11 in the first place.

The Article 50 problem nobody's SOP covers

Here is a live obligation that most promotional review procedures do not mention at all. Article 50 of the AI Act requires providers of AI systems generating synthetic audio, image, video or text to ensure the outputs are marked in a machine-readable format and detectable as artificially generated. Those duties applied from 2 August 2026. Regulation (EU) 2026/1744 made no substantive change to them, but did add a transitional grace period: providers of generative systems already on the market before 2 August 2026 have until 2 December 2026 to meet the machine-readable marking obligation.

Agencies use generative tools to produce imagery. Some of that imagery ends up in promotional assets. The question of whether an AI-generated element in a certified asset carries the required provenance marking, and who verifies it before certification, is a new item on the checklist that almost no company has added. It is not a theoretical exposure either: FDA's September 2025 announcement named AI-generated health content explicitly as within its expanded social media oversight.

The mechanism point is that this check is decidable — provenance metadata either exists in the file or it does not — which makes it precisely the kind of thing an agent should do, and precisely the kind of thing no MLR agent's marketing material currently claims to do. Ask about it.

Where the human still signs

The certification obligation has not moved, and the code that governs it does not contemplate machines at all. I text-extracted the 2024 ABPI Code PDF from the PMCPA website on 30 August 2026 and searched the resulting 4,709 lines: "artificial intelligence" occurs zero times, "machine learning" occurs zero times, and the standalone token "AI" occurs zero times. Whatever an agent does, it does inside an obligation structure written for people.

Clause 8.1 is unambiguous. Promotional material must not be issued unless its final form, to which no subsequent amendments will be made, has been certified by one person on behalf of the company. That person must be a registered medical practitioner or a pharmacist registered in the UK, or a UK-registered dentist for a product for dental use only. That person must not be the person responsible for developing or drawing up the material.

Clause 8.5 sets out what the certificate says: that the signatory has examined the final form of the material to ensure that in their belief it is in accordance with the requirements of the relevant advertising regulations and the Code, not inconsistent with the marketing authorisation and the summary of product characteristics, and a fair and truthful presentation of the facts about the medicine. It also carries an obligation that gets forgotten in automation business cases — material still in use must be recertified at intervals of no more than two years.

Clause 8.6 sets the record: companies must preserve certificates, the material in the form certified, information indicating the persons to whom it was addressed, the method of dissemination and the date of first dissemination, for not less than three years after the final use of the material, and produce them on request from the MHRA or the PMCPA.

Three consequences follow for anyone designing an agentic workflow. The signatory's belief is personal and cannot be delegated to a system, so an agent's output is evidence the signatory considers, not a substitute for their examination. The two-year recertification clock in Clause 8.5 means your approved corpus turns over whether or not you run an agent over it — which is the natural, defensible occasion to run one. And the Clause 8.6 record now has to include what the agent said, because a signatory who certified a piece the agent flagged has made a decision that will need explaining three years later.

That last point is the one to design for first. If the agent raises 1.3 findings per page and the signatory accepts most of them as non-issues, the acceptance rationale is the artefact that matters in an inspection, and it is the artefact that no workflow produces by default.

What this means in practice

Start by refusing to treat the vendor number as a finding about your company. It is a finding about somebody's corpus under somebody's definition of a gap. Convert it into a number you own by building a small adjudication set: sample 200 to 400 pages from live approved material, stratified by asset type and market, run the agent, and have two signatories independently classify every finding as a real defect, an accepted judgement or a false positive, with a third resolving disagreements. That exercise costs perhaps three to five signatory-days. It gives you a precision figure and a defect rate for your own library, and it is the only credible basis for a business case. Do it before procurement, not after.

Then fix the input rather than the reviewer. The library refactor is the work that pays regardless of which agent you eventually run, because every agent reads the same corpus and the MEDAL results say the corpus sets the ceiling. Practically: atomic claim records with stable identifiers, locator-level references rather than document-level ones, an explicit drawn-up or last-revised date on every record as Article 92(1) requires, market applicability, evidence type, and an expiry tied to the two-year recertification clock. Measure progress with the Clause 18.2 test — how long does it take to produce substantiation for an arbitrary claim.

Budget for the retrospective decision before the pilot. Write down, and have signed by the person who owns promotional compliance, what happens when the agent finds a defect in material currently in the field: who assesses materiality, on what timescale, what triggers withdrawal, and how the decision is recorded. Companies that skip this end up running the pilot on a synthetic sandbox to avoid the question, which produces a result that tells them nothing.

Rewrite the reviewer's job rather than shortening it. The value in an agent is not that a signatory reads faster. It is that a signatory receives a piece with every claim pre-linked to its locator, every figure checked against its source, every reference date checked against its expiry, and a short list of things that need a human judgement. That is a different document arriving in a different queue, and it needs a workflow change, a training change and a new set of quality metrics. Vendor-published benchmarks give a manual state of roughly a 21-day cycle and around three review rounds at $2,500 to $5,000 per asset, against AI-augmented pilots at three to five days and 1.2 rounds — figures published by Indegene and worth treating as directional rather than as evidence, since they are vendor-collected and not independently adjudicated.

Watch the denominator on the whole exercise. Veeva Pulse, drawing on more than 600 million HCP interactions a year from a large share of commercial biopharma field teams, reported on 29 May 2025 that nearly 80% of approved content is rarely or never used. Halving the review time on assets nobody deploys improves a metric and not a business, which is why the content utilisation question belongs in the same conversation as the review automation question — the case for that is set out in more detail in the analysis of how much approved pharma content never reaches the field.

Finally, the vendor questions worth asking, in order. What is the definition of a compliance gap in your yield figure, and who adjudicated it? What is the precision, not the recall, on a customer-labelled set? What structure do you require of the claims library, and what happens to accuracy when the source is a highlighted PDF rather than a structured record? Does the system check provenance marking on AI-generated elements under Article 50? What does the audit record look like when a signatory overrides a finding? A vendor with good answers to the last two is thinking about the same problem you are. A vendor who answers the first with a case study rather than a method has told you the number is marketing.

None of this depends on picking a winner. As of 30 August 2026 I could find no independently measured, customer-published outcome for any agentic MLR deployment — the numbers in circulation are all vendor-collected, which is a reason to build your own measurement rather than a reason to wait. The corpus work, the adjudication set and the certification record are agent-independent. They are also the only part of this that a competitor cannot buy from the same catalogue you did.

Questions people ask about this

What does agentic MLR review actually automate?
It automates the mechanically decidable checks: whether a claim appears in the approved label, whether a cited reference supports the sentence it is attached to, whether important safety information is present, whether a figure is faithfully reproduced from its source. It does not automate net impression, which is the judgement most FDA untitled letters in 2026 have turned on, and it cannot hold a certificate.
Is an MLR review agent high-risk under the EU AI Act?
Promotional review is not listed in Annex III, so an MLR agent is generally not a high-risk system. What already applies is Article 4 AI literacy and, since 2 August 2026, the Article 50 transparency duties, including machine-readable marking of AI-generated synthetic content. Regulation (EU) 2026/1744 deferred Annex III high-risk obligations to 2 December 2027 but left Article 50 in force.
Who has to certify promotional material in the UK?
Under Clause 8.1 of the 2024 ABPI Code, one named person on behalf of the company, who must be a registered medical practitioner or a UK-registered pharmacist, or a UK-registered dentist for dental-only products. That person must not be the person who developed the material. Clause 8.5 requires the certificate to state that the final form is not inconsistent with the marketing authorisation and summary of product characteristics.
Does the FDA require broadcast adverts to carry the full brief summary now?
No. As of 30 August 2026 the adequate provision option in 21 CFR 202.1 remains available. FDA and HHS announced on 9 September 2025 that they intend to rescind it through notice-and-comment rulemaking, listed as RIN 0910-AJ14, with a proposed rule expected in December 2026. Until a final rule is published and effective, the existing option stands.
How long must promotional certification records be kept?
Clause 8.6 of the 2024 ABPI Code requires companies to preserve certificates and the accompanying information for not less than three years after the final use of the material, and to produce them on request from the MHRA or the PMCPA. Clause 8.5 separately requires material still in use to be recertified at intervals of no more than two years.