Table of Contents
Main article link
The main article is here.
Glossary
Each entry gives a short description, the reference(s) in this article that use it, and a link for further reading outside the paper under review.
ADME / ADMET — absorption, distribution, metabolism, excretion (and toxicity). The set of properties that determine what a body does to a compound, as opposed to what the compound does to a target. Bender et al. treat them as the dividing line between ligand design and drug design: a ligand programme optimising only for on-target activity need not consider them, whereas a drug must be compatible with a human dose and regimen. See the paper's Box 1. Refs [1]. Further reading: OpenADMET, the ARPA-H-funded data-generation consortium named in Rule 12.
ALS — amyotrophic lateral sclerosis. A progressive neurodegenerative disease characterised by selective loss of motor neurons. Bender et al. use it twice: as an example of cell-type-specific disease vulnerability, and as their cited success story for human iPSC-derived models, in which patient-derived motor neurons were used to identify disease-relevant phenotypes and test ropinirole and bosutinib. Refs [1]. Further reading: NINDS ALS information page.
agentic system. A system that wraps a language model in a loop with tools, memory and feedback, so that it plans, acts, observes and revises rather than emitting one answer. Qi et al.'s inclusion criteria are multi-step reasoning, autonomous task planning, persistent state or memory, and closed-loop interaction with tools, databases or experimental environments; they explicitly exclude single-turn inference, prompt-based pipelines and one-shot retrieval-augmented generation. Refs [4], [3]. Further reading: the author's own series on agentic programming in practice [12], [13], [14].
AI — artificial intelligence. Used here in the paper's own broad sense: everything from classical statistical learning through deep learning to generative models, applied anywhere from molecular structures to trial design. Refs [1].
AlphaFold. Deep-learning system for protein structure prediction, trained on roughly fifty years of experimentally determined structures in the PDB. Bender et al. use it as the canonical example of AI succeeding where a genuine ground truth exists, and as a caution that folding accuracy is model validation, not drug discovery process validation. Refs [1]. Further reading: Jumper et al. 2021, Nature; AlphaFold Protein Structure Database.
applicability domain. The region of input space in which a model's predictions are supported by its training data. Outside it, predictions are extrapolation. The paper's Fig. 3b shows four ADME models with almost disjoint applicability domains. Refs [1]. Further reading: Applicability domain (Wikipedia).
ARPA-H — Advanced Research Projects Agency for Health. A United States federal funding agency modelled on DARPA, created to support high-risk, high-reward biomedical projects. Named by Bender et al. as the funder, at US$30 million, of the OpenADMET consortium. Refs [1]. Further reading: arpa-h.gov.
AUC — area under the (receiver operating characteristic) curve. A threshold-free summary of a binary classifier's ability to separate two classes. Bender et al.'s objection is not that it is wrong but that it is generic: two models with the same AUC can differ severalfold in the setting that matters (Rule 6). Refs [1]. Further reading: Receiver operating characteristic (Wikipedia).
BBB — blood–brain barrier. The endothelial barrier restricting passage of compounds from blood into brain. Predicting permeation across it is one of the four ADME endpoints whose datasets the paper's Fig. 3b compares. Refs [1]. Further reading: Blood–brain barrier (Wikipedia).
biomarker. A measurable indicator used to select or monitor patients. In the paper's Fig. 1b, biomarker-based cohort selection roughly halves the capitalized cost per successful launch, entirely through raised clinical-phase success rates. Refs [1]. Further reading: the FDA–NIH BEST (Biomarkers, EndpointS, and other Tools) glossary.
capitalized cost per successful launch. The total cost of one approved drug including the cost of capital and the cost of all the failures along the way. It is the denominator that makes phase II success rates dominate the return calculation in Rule 2. Refs [1]. Further reading: DiMasi et al. 2016, J. Health Econ..
CACHE — Critical Assessment of Computational Hit-finding Experiments. A blind, prospective community challenge in which computationally predicted hits are synthesised and assayed after submission, so that the ground truth does not exist when the prediction is made. Refs [1]. Further reading: cache-challenge.org.
CASP — Critical Assessment of techniques for protein Structure Prediction. The biennial blind assessment of protein structure predictors against experimentally determined structures withheld until after predictions are submitted; the venue at which AlphaFold 2 made its name. Refs [1]. Further reading: predictioncenter.org.
Cell Painting. A multiplexed fluorescent assay staining several organelles simultaneously, so that each imaged cell yields a high-dimensional morphological profile. See Box A in Part 1. Refs [1], [5], [9]. Further reading: Bray et al. 2016, Nature Protocols (the protocol); the JUMP Cell Painting Consortium.
CYP — cytochrome P450. The enzyme superfamily responsible for most oxidative drug metabolism and for activating many prodrugs. Inter-individual variation in CYP expression — genetic, or induced by smoking — is one of Bender et al.'s worked examples of why the same dose of the same drug produces different outcomes in different patients. Refs [1]. Further reading: Cytochrome P450 (Wikipedia).
ChEMBL. A manually curated database of bioactive molecules with drug-like properties, compiled from the medicinal chemistry literature. One of the two sources for the ">10⁶ bioactive ligands" figure in the paper's Box 1. Refs [1]. Further reading: ebi.ac.uk/chembl.
conditional (of biological data). The property that the result of a perturbation depends on a large number of other variables, so that the "same" experiment gives different answers in different contexts. Bender et al. name genetics, dose, comorbidity, diet, gut microbiome and smoking-induced CYP expression among them. Refs [1].
CRISPR. Genome-editing system used here for arrayed or pooled knockout screens, which supply the perturbation half of a phenotypic map (for example the 17,063 gene knockouts in Recursion's RxRx3, or insitro's pooled CellPaint-POSH screens). Refs [1], [5], [9]. Further reading: Jinek et al. 2012, Science.
DRUG-seq — Digital RNA with pertUrbation of Genes. A miniaturised, high-throughput RNA-sequencing method for reading the transcriptional response of cells to compound treatment at a small fraction of the cost of standard library preparation, used to group compounds by mechanism of action [2]a. Named by Bender et al. as one of the readouts a purpose-built data consortium would generate (Rule 12). Not to be confused with the ligand/drug distinction, with which it shares nothing but three letters. Refs [1], [2].
DMTA — design, make, test, analyse. The iterative industrial cycle of medicinal chemistry. Embedding AI models inside it, rather than beside it, is the paper's operational recommendation in Rule 11. Refs [1]. Further reading: Plowright et al. 2012, Drug Discovery Today.
drug. A compound that binds and engages its target and demonstrates acceptable efficacy and safety across preclinical and clinical endpoints. Contrast ligand; see the paper's Box 1. Refs [1].
endotype. A subtype of a clinical condition defined by a distinct biological mechanism rather than by symptoms. Endotyping is one of the "right patient" endpoints in the paper's Table 1. Refs [1]. Further reading: Anderson 2008, Lancet, where the term was coined for asthma.
epistemic opacity. The condition of not understanding, at a theoretical level, the process connecting an experimental input to its output — so that one cannot say in advance when a learned relationship will hold. Refs [1]. Further reading: Humphreys 2009, Synthese, which introduced the term for computational science.
Eroom's law. Moore's law spelled backwards: the observation that the inflation-adjusted price of one approved drug has climbed roughly exponentially for decades. See Box B in Part 1. Refs [1]. Further reading: Scannell et al. 2012, Nature Reviews Drug Discovery, the paper that named it.
Goodhart's law. Usually stated as "when a measure becomes a target, it ceases to be a good measure". Invoked by the paper against benchmark-driven progress in drug discovery. Refs [1]. Further reading: Goodhart's law (Wikipedia).
IL-1β — interleukin-1β. A pro-inflammatory cytokine released on inflammasome activation, and the conventional single-number readout of that process. Bender et al. use it as the contrast case for high-dimensional phenotyping: classifying inflammasome-perturbed cell states rather than measuring one secreted cytokine. Refs [1]. Further reading: UniProt P01584.
iPSC — induced pluripotent stem cell. An adult somatic cell reprogrammed to a pluripotent state, from which disease-relevant human cell types (motor neurons, dopaminergic neurons) can be differentiated. The paper's cited success is iPSC-derived ALS motor neurons used to identify phenotypes and test ropinirole and bosutinib. Refs [1]. Further reading: Takahashi & Yamanaka 2006, Cell.
ligand. A molecule that binds a target in a biochemical system or assay, judged on potency endpoints such as IC50. See the paper's Box 1. Refs [1].
LLM — large language model. A model trained to predict the next token in text. Bender et al. stress that such models process statistical patterns in language rather than reasoning (Rule 7). Contrast agentic system. Refs [1], [4]. Further reading: Large language model (Wikipedia).
local (of chemical space). The property that structurally small changes can produce large, non-smooth changes in activity, so that a model's behaviour does not interpolate reliably between chemical series. Refs [1].
ML — machine learning. The subset of AI concerned with fitting models to data. Used in this document interchangeably with the statistical-learning half of AI. Refs [1].
Mendelian randomization. Use of germline genetic variants as instrumental variables to estimate the causal effect of an exposure on an outcome, exploiting random allocation of alleles at conception. Bender et al. list it among the ways of injecting prior biological knowledge into otherwise data-driven inference. Refs [1]. Further reading: Sanderson et al. 2022, Nature Reviews Methods Primers.
neuro-symbolic. An architecture in which explicit symbolic structures — ontologies, typed schemas, deterministic verifiers — guide and constrain neural computation. Qi et al. describe such agents as coupling data-driven perception with structured knowledge representation. Refs [4], [14]. Further reading: Garcez & Lamb, Neurosymbolic AI: the 3rd wave.
organoid. A three-dimensional, self-organising in vitro tissue model. Cited both as a recommended data-generation system (Rule 9, Rule 11) and, in the paper's liver-spheroid example, as a system whose readouts still predict clinical risk poorly. Refs [1]. Further reading: Organoid (Wikipedia).
PDB — Protein Data Bank. The public archive of experimentally determined macromolecular structures, about fifty years deep, which supplied AlphaFold's ground truth. Refs [1]. Further reading: rcsb.org.
PK — pharmacokinetics. The time course of a drug's absorption, distribution, metabolism and excretion in a body; what determines whether a potent compound ever reaches its target at a useful concentration. Refs [1].
PubChem. A large open chemical database including bioactivity data from screening. With ChEMBL, the source of the ">10⁶ bioactive ligands" figure. Refs [1]. Further reading: pubchem.ncbi.nlm.nih.gov.
RWE — real-world evidence. Evidence derived from routinely collected health data such as electronic health records, used for adverse-event detection, biomarker discovery and identifying underdiagnosed disease. Bender et al. note that such data suffer incompleteness, inadequate ontologies and bias. Refs [1]. Further reading: FDA on real-world evidence.
SAR — structure–activity relationship. The relationship between a change in chemical structure and the resulting change in measured biological activity. Refs [1]. Further reading: Structure–activity relationship (Wikipedia).
SAMPL — Statistical Assessment of the Modelling of Proteins and Ligands. A blind, prospective community challenge for protein–ligand modelling, in which predictions of properties such as binding affinity and solvation are submitted before the measurements are released. Refs [1]. Further reading: samplchallenges.org.
Bemis–Murcko scaffold. The ring systems and connecting linkers of a molecule, with side chains removed — a coarse structural identity used to ask whether two datasets cover the same chemistry rather than merely the same compounds. The paper's Fig. 3b measures overlap both ways. Refs [1]. Further reading: Bemis & Murcko 1996, J. Med. Chem. 39, 2887–2893, doi:10.1021/jm9602928.
scRNA-seq — single-cell RNA sequencing. Transcriptome measurement in individual cells, giving a high-dimensional readout of cell state. Recommended by the paper as a functional readout for discriminating physiological states; also flagged as facing the problem of getting from high-dimensional data to decision variables. Refs [1], [10]. Further reading: Single-cell sequencing (Wikipedia).
SOTA — state of the art. The best published benchmark result. "SOTA-chasing" is the paper's term for optimising against it at the expense of real-world relevance. Refs [1].
spatial transcriptomics. Transcriptome measurement that preserves the position of cells within a tissue, so that cell state can be related to tissue architecture and neighbourhood. Refs [1], [10], [11]. Further reading: Nature Methods Method of the Year 2020.
patient stratification. Assigning patients to trial arms or treatments on the basis of a predictive marker, so that the treated population is enriched for likely responders. The lever quantified in Rule 2. Refs [1].
underspecification. The machine-learning phenomenon in which many models fit the training distribution equally well but behave differently under deployment shift, so that benchmark improvements do not translate into downstream gains. Refs [1]. Further reading: D'Amour et al., Underspecification presents challenges for credibility in modern machine learning.
virtual cell. A computational system that connects a perturbation — disease or drug — to a phenotype-level outcome. Treated by the paper as a long-term goal rather than a current capability; introduced in §1.2 and used in Rule 9 and Rule 12. Refs [1], [11]. Further reading: Bunne et al. 2024, Cell, "How to build the virtual cell with artificial intelligence".
References
Bender, A., Thomas, M. C., Scannell, J. W., Shaywitz, D. A., Ghiandoni, G. M., Greener, J. G., Pruteanu, L.-L., Jacobson, R. D., Handa, K., Hirano, M., Seal, S., Mahale, M., Schmidt, M. F., Ahfeldt, T., Grisoni, F. & Cortes-Ciriano, I. Artificial intelligence in drug discovery — what it is, where we stand and the path forward. Nature Reviews Drug Discovery (2026). doi:10.1038/s41573-026-01496-2. Open access: https://www.nature.com/articles/s41573-026-01496-2 (tables are served at
/tables/1,/tables/2and/tables/3).Ye, C., Ho, D. J., Neri, M., Yang, C., Kulkarni, T., Randhawa, R., Henault, M., Mostacci, N., Farmer, P., Renner, S., Ihry, R., Mansur, L., Gubser Keller, C., McAllister, G., Hild, M., Jenkins, J. & Kaykas, A. DRUG-seq for miniaturized high-throughput transcriptome profiling in drug discovery. Nature Communications 9, 4307 (2018). https://pmc.ncbi.nlm.nih.gov/articles/PMC6192987/
Huynh, D. L., Seal, S., AIA4S Consortium, Reid, D., Carpenter, A. E., Bender, A. & Spjuth, O. AI agents in drug discovery: applications and case studies. Drug Discovery Today 31(3), 104650 (2026); article type KEYNOTE; open access under CC BY. doi:10.1016/j.drudis.2026.104650 — https://doi.org/10.1016/j.drudis.2026.104650. Preprint version: arXiv:2510.27130 (2025), under the earlier author order Seal, S., Huynh, D. L. et al. Cited by Bender et al. [1] as their reference 200.
Qi, C., Wang, W., Jiang, S., Liu, Q., Song, X., Fang, H. & Wei, Z. Artificial Intelligence agents for biological research: a survey. Briefings in Bioinformatics 27(1), bbag075 (2026). https://academic.oup.com/bib/article/27/1/bbag075/8499367
Recursion Pharmaceuticals. RxRx3: Phenomics Map of Biology. https://www.rxrx.ai/rxrx3 (accessed 2026-09-18).
Artificial Intelligence in Small-Molecule Drug Discovery: A Critical Review of Methods, Applications, and Real-World Outcomes. Pharmaceuticals 18, 1271 (2025). https://pmc.ncbi.nlm.nih.gov/articles/PMC12472608/
Minikel, E. V., Painter, J. L., Dong, C. C. & Nelson, M. R. Refining the impact of genetic evidence on clinical success. Nature 629, 624–629 (2024). https://pmc.ncbi.nlm.nih.gov/articles/PMC11096124/
Markowetz, F. All models are wrong and yours are useless: making clinical prediction models impactful for patients. npj Precision Oncology 8, 54 (2024). https://pmc.ncbi.nlm.nih.gov/articles/PMC10901807/
insitro. insitro Validates AI-Enabled POSH Platform in Nature Communications, Bridging Critical Gap in Drug Discovery. Press release, 16 December 2025. https://www.insitro.com/news/insitro-validates-ai-enabled-posh-platform-in-nature-communications-bridging-critical-gap-in-drug-discovery/ (accessed 2026-09-18).
Relation Therapeutics. Our science. https://www.relationrx.com/science (accessed 2026-09-18).
Noetik. Applications. https://www.noetik.ai/applications (accessed 2026-09-18).
Donaldson, I. Agentic programming in practice, part 1. blog.recursivehuman.com (accessed 2026-09-18). https://blog.recursivehuman.com/p/agentic-programming-in-practice-part-1
Donaldson, I. Agentic programming in practice, part 2. blog.recursivehuman.com (accessed 2026-09-18). https://blog.recursivehuman.com/p/agentic-programming-in-practice-part-2
Donaldson, I. Agentic programming in practice, part 3. blog.recursivehuman.com (accessed 2026-09-18). https://blog.recursivehuman.com/p/agentic-programming-in-practice-part-3 ---
Provenance and verification
What changed in this version. Earlier drafts wove verbatim quotations into the prose and linked each into a page of a locally cached PDF. Those links were never going to work for a reader, and the quotations made the article read like a concordance. This version paraphrases; every claim taken from a source carries a hyperlinked [N]; and where a reviewer might want to cross-examine a claim, the citation is followed by superscript letters, each a link that opens the source's open-access page scrolled to the supporting passage. The per-quotation verification tables that used to live here have therefore collapsed into the References: the footnotes are the verification rows. What remains here is the method and the scripts' output.
Provenance footnotes: method. Every footnote is a row of document_11_figs/data/provenance_table.py — tag, reference number, locator, the exact supporting passage, and a note giving section and PDF page. For a web locator the #:~:text= fragment is built from the passage by the debate skill's link_reference.py (never by hand), so that a Chromium browser scrolls to and highlights the passage. verify_provenance.py fetches every locator, reduces the page to plain text exactly as a browser's fragment search sees it, and requires the passage to be a verbatim substring — no normalisation and no fallback, because the HTML and the PDF text layer differ in hyphenation, punctuation and where citation numerals sit, and a fragment that is off by one character simply fails to highlight. Where the passage runs into an inline reference numeral on the page, the footnoted span was shortened to stop before it, and the note says so. Two sources could not be fetched by script: the Huynh/Seal/Bender keynote sits behind a ScienceDirect page that answers HTTP 403 to non-browser clients, and Qi et al. behind a Cloudflare challenge on academic.oup.com. Their five footnotes name a PDF page instead, unlinked, and are checked against that page's text layer with the same four normalisation rules earlier drafts documented (soft hyphens closed, superscript numerals stripped, kerning splits after T/V/W/Y/P closed, space before punctuation removed). The script also checks that every [^…] marker in the prose is defined exactly once and that no table row is unused.
Provenance footnotes: result.
$ python3 verify_provenance.py
structure: 88 tags used, 88 definitions, 0 problems
88 verified, 0 failed (86 web fragments, 2 PDF page anchors)
No verbatim copying: method. document_11_figs/data/verify_no_verbatim.py takes this document, removes everything that is supposed to reproduce the paper's wording — text inside quotation marks, fenced blocks and blockquotes, the reference list and its footnote definitions, this appendix's script output — and everything that is not prose (headings, HTML anchors, URLs, citation labels), lower-cases and tokenises what is left, and reports every run of eight or more consecutive words that also occurs in background/bender2026_fulltext.txt, the text extraction of the paper. Eight is the threshold the author set; the script takes it as a parameter so the Critic can lower it to look for close paraphrase.
No verbatim copying: result.
$ python3 verify_no_verbatim.py
document_11.md: 8530 prose tokens compared against 22600 source tokens; threshold 8 words; blockquotes excluded
line 213 9 words: that generating data is not the bottleneck in drug
1 shared run(s) of >= 8 consecutive words found
$ python3 verify_no_verbatim.py --min-words 6
document_11.md: 8530 prose tokens compared against 22600 source tokens; threshold 6 words; blockquotes excluded
line 42 6 words: artificial intelligence ai in drug discovery
line 108 6 words: for ai in drug discovery are
line 167 6 words: to the use of ai in
line 203 11 words: space is by definition unknown at the time a model is
line 209 6 words: a measure becomes a target it
line 213 9 words: that generating data is not the bottleneck in drug
line 247 6 words: not be relevant if the concentration
7 shared run(s) of >= 6 consecutive words found
The one run at the eight-word threshold is in Rule 5 and is the author's own sentence — "the important but often forgotten maxim that generating data is not the bottleneck in drug discovery" — which happens to share nine words with the paper's "generating data is not the bottleneck in drug discovery — it is 'reduction to practice'". Rules 1, 2, 4, 5 and 6 are the author's text and were not rewritten by the Writer; the overlap is flagged for the author's decision rather than paraphrased on his behalf. The six-word near-misses are, apart from acronym expansions and the phrase "AI in drug discovery", also in the author's own paragraphs (Rules 2, 4, 5 and 6) and are reported for the same reason.
Citation mechanics. fix_citations.py --in-place was run before the audit:
$ python3 fix_citations.py document_11.md --in-place
wrote document_11.md
entries : 14 (0 renumbered)
in-text citations : 0 linked (0 renumbered)
quote links paired : 0 (0 already carried a reference)
audit_document.py --skip quote-without-reference (the skipped check enforces the superseded quote-link style):
$ python3 audit_document.py document_11.md --skip quote-without-reference --summary-only
# audit_document — `document_11.md`
| Check | high | medium | low |
|---|---:|---:|---:|
| `undefined-term` | 0 | 5 | 1 |
| `late-introduction` | 0 | 8 | 0 |
| `register` | 20 | 5 | 1 |
| `unsourced-attribution` | 0 | 1 | 0 |
**41 candidate(s).** Confirm each before acting on it.
Which sources have been read in full. Reference [1] — the paper under review — has been read end to end, as has the keynote [3]; both text extractions are in the debate area. Reference [14] was read in full in an earlier round after a draft had misdescribed it. References [2], [4]–[13] were read at the level of the passages footnoted from them, and claims made about those sources should be read as claims about those passages. The company pages [5], [9]–[11] are marketing material; they are cited only for what each company says its own platform does, which is the claim Rule 7 makes.
A limitation of quotation-checking, kept on the record. Confirming that a passage appears verbatim on a page establishes wording fidelity. It does not establish content fidelity: earlier drafts of this debate twice built a wrong claim around a character-perfect quotation, once about the author's own blog series and once about the keynote, and in both cases the remedy was to read the source in full rather than to tighten the string comparison. The material carrying those corrections has moved to the follow-on article's seed file; the lesson is recorded here so that it travels with this one.
Figures. None. Earlier drafts carried six; the author removed the four that restated a two-number contrast from the paper, and the two that concerned the agentic-AI field went with the material cut for the follow-on article. document_11_figs/ therefore holds only the data/ scripts above and their manifest.
