Article

From pilot to protagonist

Pharma is handing AI the decisions, not just the slides. Revisited seven weeks on, with a finding that moves the argument: the machine's characteristic error is no longer inventing sources but quietly leaving things out -- which is harder to catch, and strengthens the case rather than softening it.

agentic AI in drug development, and why provenance is the moat · 2026-08-06 · 12 sources

First published 15 June 2026. Revisited 6 August 2026 against the literature published since, which changed one of its load-bearing claims. The original argument is unaltered below; what is new is marked, and the section "What we worry about now" is the reason for the revisit.

For three years the drug industry's romance with artificial intelligence was mostly choreography. Models drew protein structures, summarized trials and flattered steering committees. The work was real, but the verb was always "assist": a human decided, the machine suggested.

The new grammar

Pharma has quietly changed the verb from "suggest" to "decide" -- and the verb is the whole story.

That grammar is changing, and the change is now on the record. On June 5, 2026 Owkin and Sanofi widened a partnership that began in 2021 with a EUR 90m oncology collaboration into a five-year license for "K Pro" -- a platform Owkin bills not as a copilot but as an "AI scientist": a job title, not a feature.1Owkin. (2026, June 5). "Owkin to build AI agents as part of a multi-year K Pro collaboration with Sanofi" [Press release]. owkin.com. Five-year K Pro license; builds on the 2021 EUR 90m partnership.Primary Three months earlier, at NVIDIA's GTC, IQVIA unveiled IQVIA.ai, a unified agentic platform spanning clinical, commercial and real-world operations; the firm reports more than 150 agents in production and says 19 of the top 20 pharmaceutical companies already use them -- the twentieth, presumably, is still in procurement.2IQVIA. (2026, March 16). "IQVIA unveils IQVIA.ai, a unified agentic AI platform" [Press release]. Announced at NVIDIA GTC. iqvia.com. More than 150 agents, 19 of the top 20 pharma; clinical / commercial / real-world operations, not bench discovery.Primary Benchling's 2026 survey of around 100 biotech organizations catches the mood from the inside: "build what differentiates, buy what scales" -- make-versus-buy, with higher stakes and worse documentation.3Benchling. (2026). "2026 biotech AI report." benchling.com. Survey of about 100 organizations, US and EU, November 2025. Adoption 76 / 71 / 66 / 58%; "adoption drops where data is scattered, incomplete, hard to validate." The secondary "29-42%" band is not in the primary and is excluded.Primary

The throughline is a single, consequential shift. AI is moving from the margins of the workflow toward its center -- the place where decisions are made. And the center is a different country from the edge.

The bill for being wrong

A tool that drafts can be overruled in a meeting; an engine that decides cannot -- and the ledger on machine error does not flatter it.

This is, mostly, good news. Design-make-test-analyze cycles compress; a biomarker hypothesis that once took a quarter can take a fortnight. But a decision engine is a different animal from a drawing tool, and it imports a different risk. When a model suggests, a wrong answer costs a meeting. When a model decides -- or shapes a decision so heavily that no human re-derives it -- a wrong answer costs a program, a partnership, sometimes a patient.

Here the evidence is sobering, and it is not anecdote. A 2023 audit in the Journal of Clinical Medicine asked a large language model for nephrology references: of 610 it produced, only a fifth were authentic -- the rest were fiction with footnotes.4Suppadungsuk, S., Thongprayoon, C., Krisanapan, P., Tangpanithandee, S., Garcia Valencia, O., Miao, J., Mekraksakit, P., Kashani, K., & Cheungpasitporn, W. (2023). "Examining the validity of ChatGPT in identifying relevant nephrology literature: Findings and implications." Journal of Clinical Medicine 12(17), 5550. DOI 10.3390/jcm12175550. PMID 37685617. Of 610 references: 62% existed, 20% authentic, 31% fabricated.Primary The trouble runs deeper than citations. Kapoor and Narayanan, writing in Patterns, found data leakage corrupting machine-learning results across 294 papers in 17 scientific fields; in a worked correction, an elaborate model lost to logistic regression -- a method old enough to collect a pension.5Kapoor, S., & Narayanan, A. (2023). "Leakage and the reproducibility crisis in machine-learning-based science." Patterns 4(9), 100804. DOI 10.1016/j.patter.2023.100804. PMID 37720327. Leakage across 294 papers in 17 fields.Primary And the commercial reckoning has arrived: MIT's NANDA initiative reported in 2025 that 95% of enterprise generative-AI pilots produced no measurable effect on profit and loss; an estimated $30bn-40bn bought a great deal of enthusiasm and almost no profit.6MIT NANDA. (2025). "The GenAI divide: State of AI in business 2025." Massachusetts Institute of Technology. 95% of enterprise GenAI pilots produced no measurable P&L impact. No PMID; figure quoted from the report, not deep-linked.Secondary The machine is fast. It is not, by default, right.

Added 6 August 2026. That nephrology figure is the claim this piece would now lead with differently. It still stands as measured, but it is no longer the error worth worrying about most, and the section below explains what displaced it.

Where the machine earns its seat

Adoption tracks one thing only -- whether the answer can be checked -- which is why AI sprints through chemistry and limps through biology.

What separates the two is unglamorous: provenance -- the unbroken chain from a stated fact back to the primary source that licenses it, gradeable and auditable rather than vibe-checked. It is a discipline, not a dashboard: every claim atomic, every claim tagged by the strength of its evidence, every figure traceable to a document a skeptic can open.

The empirical pattern bears this out. Benchling's own figures show adoption is high where the ground truth is clean -- literature review (76%), protein-structure prediction (71%), scientific reporting (66%), target identification (58%) -- and drops sharply in generative design, biomarker analysis and ADME, "where data is scattered, incomplete, and hard to validate."3Benchling. (2026). "2026 biotech AI report." benchling.com. Survey of about 100 organizations, US and EU, November 2025. Adoption 76 / 71 / 66 / 58%; "adoption drops where data is scattered, incomplete, hard to validate." The secondary "29-42%" band is not in the primary and is excluded.Primary The clinic tells the same story. An analysis in Drug Discovery Today found AI-originated molecules clear Phase I at 80-90%, well above the norm, then fall to roughly 40% in Phase II -- the industry average.7Jayatunga, M. K. P., Ayers, M., Bruens, L., Jayanth, S., & Meier, C. (2024). "How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons." Drug Discovery Today 29(6), 104009. DOI 10.1016/j.drudis.2024.104009. PMID 38692505. 80-90% Phase I, about 40% Phase II.Primary AI aces the exam it can cram for and stumbles on the one nobody can: it makes drug-like molecules beautifully, and is so far no better than anyone else at the part that matters, whether the biology is real.

The fix is not to slow the agents; it is to ground them. When researchers built a multi-agent literature system tethered to live PubMed retrieval rather than free generation, citation accuracy reached 99.82% with no fabricated sources -- the same models that invented citations simply stopped inventing them.8Gorenshtein, A., Shihada, K., Sorka, M., Aran, D., & Shelly, S. (2025). "LITERAS: Biomedical literature review and citation retrieval agents." Computers in Biology and Medicine 192(Pt B), 110363. DOI 10.1016/j.compbiomed.2025.110363. PMID 40383055. Retrieval-grounded: 99.82% citation accuracy, 0% non-academic sources.Primary Tool-grounding, not raw fluency, is what makes machine output defensible.

What we worry about now

Added 6 August 2026. The failure mode moved -- and it moved in the direction that is hardest to audit.

Seven weeks is not long, but the literature published since June has done something more useful than confirm the argument above. It has changed what the argument should be afraid of.

A systematic review in the Journal of Biomedical Informatics, published 25 July 2026, pooled 27 studies of large language models extracting data for evidence synthesis.9Shankar, R., Lim, A., & Qian, X. (2026). "Performance of large language models in data extraction for evidence synthesis: A systematic review." Journal of Biomedical Informatics, published 25 July 2026, 105086. DOI 10.1016/j.jbi.2026.105086. PMID 42501879. Twenty-seven studies; accuracy 47-99.9%; categorical and string variables 74-96% versus numerical 47-88%; omissions 60-74% of errors against hallucination 0.08-6%; assistive workflow 91.0% (95% CI 90.4-91.6) versus human-only 89.0%. Evidence base searched "through December 2025" across GPT-4/4o, Claude 2-3.5, Gemini, Llama, Mistral, Qwen and DeepSeek; eligible studies evaluated extraction "against a human reference standard", i.e. with the source supplied.Primary Its headline finding is not the accuracy range, though that is worth seeing on its own -- 47-99.9%, which is less a performance figure than a warning that "how accurate is it" is the wrong question without naming the task. The finding that matters is the shape of the errors. Omissions were dominant, at 60-74% of them. Fabrication, the thing this essay led with in June, ran at 0.08-6%.

Two mistakes are available here and it is worth refusing both.

The first is to conclude that fabrication was overblown. It was not, and the two results do not contradict each other, because they did not measure the same thing. The 2023 nephrology audit asked a model for references it did not have -- generation with nothing in hand. The review's inclusion criteria required evaluation "against a human reference standard": the source document was supplied, and the model's job was to read it correctly. So the honest statement is conditional, and more useful than either headline:

source withheld   ->   the model invents
source supplied   ->   the model omits

The second mistake is to file omission as the milder problem because its cousin sounds worse. It is the more dangerous of the two, for a reason this site keeps running into in its own machinery. An invented citation fails a check that already exists: the identifier does not resolve, and something says so. An omitted claim fails no check at all. Nothing in a document announces the sentence that is not in it. There is no anomaly, because absence has no shape -- which is the same reason a missing page never trips a link checker and a report with no findings looks identical to a report where nobody looked.

One more number from the same review sharpens it for anyone reading a diligence memo. Categorical and string variables were extracted reliably, at 74-96%. Numerical data was not, at 47-88%. The numbers are the least reliable thing the machine hands you, and the numbers are what the decision runs on.

None of which is an argument for doing less of this. The same review found an assistive workflow reaching 91.0% accuracy against 89.0% for human-only extraction -- the machine and a person together beat the person alone, which is the whole case for the arrangement. And grounding continues to do what June said it would, with a limit now visible: a retrieval-grounded citation system published in JAMIA on 4 August reached 100% factual accuracy while its relevance score stopped at 77.50%.10Xie, Q., Zhang, J., Wang, Y., Huang, J., Lin, F., Weng, R.-L., He, H., Chen, Q., & Xu, H. (2026). "CiteSure: retrieval-augmented large language models for faithful biomedical citation recommendation." Journal of the American Medical Informatics Association, published 4 August 2026. DOI 10.1093/jamia/ocag122. PMID 42551844. 100% factual accuracy; highest relevance score 77.50%.Primary Every source it produced was real; relevance -- whether those real sources were the ones the question called for -- is the axis that did not close, and it is a graded score rather than a count, so the gap is a direction and not a percentage of wrong answers. Grounding solved fabrication and did not solve judgement, and no amount of it will.

The durable lesson is the movement itself. Each fix relocated the failure rather than removing it: ungrounded models invented sources, so we grounded them, and now they retrieve real sources and omit the inconvenient ones. What is worth building is therefore not a defence against the error we can currently name. It is the habit of noticing when the error has changed address -- which requires having written down, in a form you can re-run, what you checked and when.

The only defensible moat

The winners will not own the cleverest agents; they will own the shortest path from any claim to the document that proves it.

The firms that win the next phase will not be the ones with the flashiest agents. They will be the ones who can answer, instantly and defensibly, a deceptively simple question about any AI-derived claim: how do you know? As pharma hands the machine a seat at the table, the premium moves to whoever still checks its work.

Added 6 August 2026. That was a contrarian sentence in June and it is not one now, which is worth saying plainly rather than claiming as a win. Clinical Pharmacology and Therapeutics published the same argument in September, peer reviewed, from the pharmacology side: the scientist's role is moving from executor to orchestrator, and the risk is "evaluating recommendations that may be difficult to verify independently."11McCoy, Michael, & McCoy, Matthew. (2026). "From Executor to Orchestrator: The Pharmacology Scientist in the Age of Agentic AI." Clinical Pharmacology and Therapeutics, September 2026. DOI 10.1002/cpt.70380. PMID 42411454. Phase I to approval "remaining near 10%"; "each automation wave increased throughput while leaving the interpretive bottleneck intact"; "evaluating recommendations that may be difficult to verify independently"; "existing governance structures do not address the failure modes that accompany delegation of scientific judgment to autonomous systems."Primary Its sharpest line is a rebuke to fifty years of tooling -- "each automation wave increased throughput while leaving the interpretive bottleneck intact" -- and it notes that the probability of a compound entering Phase I reaching approval has stayed near 10% throughout. A review in Drug Discovery Today four days earlier reaches the same requirement from the model side: "prospective validation, rigorous human oversight and governance frameworks."12Khan, S., Samuelsson, J. G., Cai, X., Natarajan, K., & Madhavan, S. (2026). "Can agentic AI dent Eroom's Law? Large reasoning models across the drug discovery and development pipeline." Drug Discovery Today, published 4 August 2026, 104754. DOI 10.1016/j.drudis.2026.104754. PMID 42551551. Large reasoning models "couple multi-step reasoning with the ability to plan, invoke external tools and retrieve authoritative information"; deployment "requires prospective validation, rigorous human oversight and governance frameworks."Primary

Consensus changes what the claim is worth. When everyone agrees the machine needs checking, saying so stops being a differentiator and the only thing left that distinguishes anyone is whether they actually do it, in a form a sceptic can re-run. The pharmacology paper puts the gap in one sentence: existing governance "does not address the failure modes that accompany delegation of scientific judgment to autonomous systems." That is not a prediction about who wins. It is an observation that the work is unbuilt.

Appendix A -- provenance and method

How every number above was pulled, graded, and made checkable -- not asserted.

I do not ask you to trust this article; I ask you to check it. Each claim was reduced to a single statement, tied to a primary document, and graded by how strong its evidence is. The pipeline below is the audit trail -- reproducible, so a skeptic can re-run it and land on the same numbers. That is what lets you act on the read instead of taking it on faith.

StepTool / API callWhat it verified
Citation retrievalPubMed.get_article_metadata x 4 (PMIDs 37685617, 37720327, 38692505, 40383055)Each cited paper exists; title, journal, year, and the quoted figure match the abstract of record.
Primary-source capturetavily.search + jina.read_url on owkin.com, iqvia.com, benchling.comOwkin K Pro five-year license; IQVIA >150 agents / 19-of-20; Benchling adoption percentages -- read from the originating pages, not a secondary summary.
Claim atomisationbin/facts fact-engine (admit + locator + quote + as_of)Every sentence-level claim reduced to one statement with a source locator and verbatim quote; no compound or unsourced assertions admitted.
Evidence gradingProvenance grade, primary / secondaryOwkin / IQVIA / Benchling / PMIDs = primary; MIT-NANDA P&L figure = secondary (report, not peer-reviewed). Grades shown in the endnotes.
Hedge / fabrication gatedoctor.py T7 + T8 (hedge-gate, R6)Rejected the unsourced "29-42%" adoption band -- present in a secondary write-up, absent from Benchling's primary -- rather than launder an estimate as fact.

The 6 August revisit, run the same way. Listed separately rather than merged into the rows above, because a provenance appendix that quietly absorbs a second session's tool calls into the first session's table is doing the thing this article objects to.

StepTool / API callWhat it verified
Literature sweepToolUniverse CLI v1.1.11, tu run PubMed_search_articles x 4 (2026 agentic-AI / evidence-synthesis terms)What had been published since 15 June. The first query returned zero results from over-specification; the four that ran were deliberately broader.
Abstract capturetu run PubMed_get_article x 4 (PMIDs 42411454, 42501879, 42551551, 42551844)Every quoted figure and phrase read from the abstract of record, not a summary of it. Titles, journals and dates as cited in the endnotes.
Task-equivalence checkInclusion criteria of PMID 42501879, read verbatimThat the 2023 and 2026 findings measure DIFFERENT tasks -- "against a human reference standard" means the source was supplied -- so the newer number does not retire the older one.
Recency boundSearch window of PMID 42501879, read verbatimIts evidence base runs "through December 2025" over GPT-4/4o, Claude 2-3.5, Gemini, Llama, Mistral, Qwen and DeepSeek. A 2026 publication date is not a 2026 evidence base.

That third row is the one that took the work. The tempting version of this update was to swap a 2023 citation for a 2026 one and let the newer number stand, which would have read as diligence and been the opposite: a claim laundered by recency. Two studies disagreeing is only a correction if they asked the same question.

The point of this appendix is the moat. An article you can audit in five clicks is doing in public exactly what pharma now needs from its AI: an unbroken, gradeable chain from claim to primary source. The rigor is the product.

Appendix B -- reasoning, run four ways

How the conclusion was reached, run four ways -- so you can find the seam if there is one.

This piece rests on one claim: as AI moves from advising to deciding in drug research, the advantage goes to whoever can prove how a claim was reached. Richard Feynman's first rule of honest thinking was that you must not fool yourself -- and you are the easiest person to fool. The cure is independence: if you reach the same answer by several routes that do not lean on each other, the odds you fooled yourself on every one grow small. I ran this claim down four such routes.

RouteWhere it lands
Follow the evidenceAI now decides -- and invents a third of its own citations
Assume it, then checkAll three conditions hold; grounding lifts accuracy to 99.82%
The recurring patternAI races where answers are checkable, stalls where they are not
From a basic principleAn unchecked decision carries its errors straight downstream

Follow the evidence

Start from the record and trace it forward to where it leads.

Just follow the trail. Owkin and Sanofi now let a platform they call an "AI scientist" help decide which programs advance; IQVIA runs more than 150 such agents across 19 of the top 20 drugmakers; Benchling's survey finds the same shift inside the labs. So AI has crossed from making the slides to making the call -- and a bad call costs a whole program, not a meeting. Because these models still invent a real share of what they produce -- a third of the references in one audit were fabricated -- the trail ends somewhere uncomfortable: what matters is no longer having AI, but being able to check it.

Assume it, then check

Suppose the claim is true; confirm each condition it would require (working backward).

Now try to break the claim instead of build it. If "you win by proving your work" were true, three things would have to be real, and each is easy to look up. AI would have to be making real decisions -- it is. Those decisions would have to fail often enough to matter -- they do, from fabricated citations to a leakage problem across 294 studies to the 95% of corporate pilots that returned no profit. And there would have to be a fix that works -- there is: tie the model to a live literature search and its citation accuracy reaches 99.82%. The claim refuses to break.

The recurring pattern

One regularity shows up across many independent cases (induction).

Step back and watch the same thing happen again and again. Where the truth is easy to check, AI is everywhere -- combing the literature, predicting protein shapes. Where the data is messy and hard to validate, trust collapses. AI-designed molecules sail through the early trials that test chemistry, which you can measure, then fall back toward the average in the later trials that test whether the drug works, which you cannot fake. One pattern explains every case: AI speeds up exactly where its answers can be checked. This is induction -- a strong pattern, not an ironclad law, so a genuinely new field could defy it -- but it holds across very different ones.

From a basic principle

From a premise almost no one rejects, the conclusion follows of necessity (deduction).

Finally, reason it out from something obvious. If a machine decides and no human re-derives the result, its mistakes pass straight into the outcome -- no one disputes that. I have also shown that ungrounded models err at a real, measured rate. Put those two facts together and the conclusion is not optional; it is forced. An unchecked AI making real decisions will push real errors into real drug programs. So checking the work -- grounding it in sources, keeping the trail -- is not polish; it is the requirement.

Four roads, one destination. Any single road you might doubt -- perhaps I picked the evidence, perhaps the pattern is a fluke. But it is hard to fool yourself the same way four independent times. That convergence is the reason to trust the answer, and it is why I show every step: a conclusion you can reach four ways, each tied to a document you can open yourself, is one you can act on.

Footnotes

  1. Owkin. (2026, June 5). "Owkin to build AI agents as part of a multi-year K Pro collaboration with Sanofi" [Press release]. owkin.com. Five-year K Pro license; builds on the 2021 EUR 90m partnership. (Primary)

  2. IQVIA. (2026, March 16). "IQVIA unveils IQVIA.ai, a unified agentic AI platform" [Press release]. Announced at NVIDIA GTC. iqvia.com. More than 150 agents, 19 of the top 20 pharma; clinical / commercial / real-world operations, not bench discovery. (Primary)

  3. Benchling. (2026). "2026 biotech AI report." benchling.com. Survey of about 100 organizations, US and EU, November 2025. Adoption 76 / 71 / 66 / 58%; "adoption drops where data is scattered, incomplete, hard to validate." The secondary "29-42%" band is not in the primary and is excluded. (Primary) 2

  4. Suppadungsuk, S., Thongprayoon, C., Krisanapan, P., Tangpanithandee, S., Garcia Valencia, O., Miao, J., Mekraksakit, P., Kashani, K., & Cheungpasitporn, W. (2023). "Examining the validity of ChatGPT in identifying relevant nephrology literature: Findings and implications." Journal of Clinical Medicine 12(17), 5550. DOI 10.3390/jcm12175550. PMID 37685617. Of 610 references: 62% existed, 20% authentic, 31% fabricated. (Primary)

  5. Kapoor, S., & Narayanan, A. (2023). "Leakage and the reproducibility crisis in machine-learning-based science." Patterns 4(9), 100804. DOI 10.1016/j.patter.2023.100804. PMID 37720327. Leakage across 294 papers in 17 fields. (Primary)

  6. MIT NANDA. (2025). "The GenAI divide: State of AI in business 2025." Massachusetts Institute of Technology. 95% of enterprise GenAI pilots produced no measurable P&L impact. No PMID; figure quoted from the report, not deep-linked. (Secondary)

  7. Jayatunga, M. K. P., Ayers, M., Bruens, L., Jayanth, S., & Meier, C. (2024). "How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons." Drug Discovery Today 29(6), 104009. DOI 10.1016/j.drudis.2024.104009. PMID 38692505. 80-90% Phase I, about 40% Phase II. (Primary)

  8. Gorenshtein, A., Shihada, K., Sorka, M., Aran, D., & Shelly, S. (2025). "LITERAS: Biomedical literature review and citation retrieval agents." Computers in Biology and Medicine 192(Pt B), 110363. DOI 10.1016/j.compbiomed.2025.110363. PMID 40383055. Retrieval-grounded: 99.82% citation accuracy, 0% non-academic sources. (Primary)

  9. Shankar, R., Lim, A., & Qian, X. (2026). "Performance of large language models in data extraction for evidence synthesis: A systematic review." Journal of Biomedical Informatics, published 25 July 2026, 105086. DOI 10.1016/j.jbi.2026.105086. PMID 42501879. Twenty-seven studies; accuracy 47-99.9%; categorical and string variables 74-96% versus numerical 47-88%; omissions 60-74% of errors against hallucination 0.08-6%; assistive workflow 91.0% (95% CI 90.4-91.6) versus human-only 89.0%. Evidence base searched "through December 2025" across GPT-4/4o, Claude 2-3.5, Gemini, Llama, Mistral, Qwen and DeepSeek; eligible studies evaluated extraction "against a human reference standard", i.e. with the source supplied. (Primary)

  10. Xie, Q., Zhang, J., Wang, Y., Huang, J., Lin, F., Weng, R.-L., He, H., Chen, Q., & Xu, H. (2026). "CiteSure: retrieval-augmented large language models for faithful biomedical citation recommendation." Journal of the American Medical Informatics Association, published 4 August 2026. DOI 10.1093/jamia/ocag122. PMID 42551844. 100% factual accuracy; highest relevance score 77.50%. (Primary)

  11. McCoy, Michael, & McCoy, Matthew. (2026). "From Executor to Orchestrator: The Pharmacology Scientist in the Age of Agentic AI." Clinical Pharmacology and Therapeutics, September 2026. DOI 10.1002/cpt.70380. PMID 42411454. Phase I to approval "remaining near 10%"; "each automation wave increased throughput while leaving the interpretive bottleneck intact"; "evaluating recommendations that may be difficult to verify independently"; "existing governance structures do not address the failure modes that accompany delegation of scientific judgment to autonomous systems." (Primary)

  12. Khan, S., Samuelsson, J. G., Cai, X., Natarajan, K., & Madhavan, S. (2026). "Can agentic AI dent Eroom's Law? Large reasoning models across the drug discovery and development pipeline." Drug Discovery Today, published 4 August 2026, 104754. DOI 10.1016/j.drudis.2026.104754. PMID 42551551. Large reasoning models "couple multi-step reasoning with the ability to plan, invoke external tools and retrieve authoritative information"; deployment "requires prospective validation, rigorous human oversight and governance frameworks." (Primary)