Table of Contents
Abstract
This is a review of a long, careful and somewhat deflationary Perspective on artificial intelligence (AI) in drug discovery. The authors take the position that it is too early to judge whether AI will be effective in drug discovery and then go on to suggest that if it is ever to be shown effective, the field must change its focus to validate processes in drug discovery rather than just the models that they operate around.
1. Introduction
1.1 What this is, and what it isn't
You might be forgiven for thinking this is another paper about how the latest razzle-dazzle AI machines are going to automagically create the greatest drugs ever, faster than we can say initial public offering. It isn't. Bender et al. [1] take a much more sober approach to the topic, and a much more expansive view of both AI and drug development.
For one thing, they correctly and refreshingly scope AI to include everything from classical machine learning (ML) to generative models, working on anything from molecular structure data to clinical trial design — something easy to forget in these heady days of LLM innovation. For another, they scope drug discovery well past target and ligand discovery: their Table 1 tracks each modellable endpoint through to the phase in which its impact actually shows up; for example, selection of targets with genetic support in the target-selection phase can give a 2.6-fold increase in the likelihood that they will be successful in the clinic during phases II and III [1]ab.
This scoping decision has consequences for both halves of the title. Most ligands are not drugs (the paper's Box 1 is devoted to the difference), and not all problems that can be attacked with AI are the problems that need solving. The authors' central recommendation is that benchmarking in this field has to graduate from model validation (does it score well on a test set?) to process validation (does it improve a real decision?). Everything else in the paper is, more or less, an argument for why that distinction is consequential.
The paper is long, but worth reading for its broad insight across two fields, the referenced research behind them, and an honest accounting of human dynamics that you don't normally encounter in a research journal.
Below I try to compress the recommendations and advice into the style of a "10 simple rules" article. Sorry — it's more complicated than 10. It came to fifteen.
1.2 The concepts
Two distinctions carry most of the weight in what follows, so they get a paragraph each here. Everything else the paper leans on — ADME, AUC, applicability domains, the public bioactivity databases, the blind benchmarking challenges, the experimental systems behind the later rules — is defined in the Glossary of the companion Appendices post.
The two data domains. Biological data means data from cellular and in vivo disease models, pharmacology, and clinical trials. Chemical data means compound structures and the properties measured on them, including structure–activity relationships (SAR) — how a change to a molecule changes what it does. The paper's claim is that both behave badly for AI, but for different reasons: biological data are conditional and epistemically opaque, chemical data are local and biased. One of the paper's recurring examples of conditionality is drug metabolism by cytochrome P450 (CYP) enzymes, whose activity varies between people and is induced by smoking.
Ligands and drugs. A ligand binds a target in an assay. A drug additionally has to be safe and efficacious in a person. The gap between them is the subject of the paper's Box 1 and of Rule 2.
1.3 The strongest version of the case this paper argues against
Advocates for AI in drug discovery are plentiful. Here is their case, in its best form, with sources. This paper doesn't dismiss that case, but rather advises caution on how early successes are interpreted and what's needed going forward.
The technical capability is real and has moved fast. Protein structure prediction went from an open problem to a solved-enough one; the paper itself concedes that diffusion-based structure and design models have attracted serious money — its example is Xaira, whose first funding round was US$1 billion [1]a. Phenotypic screening at industrial scale is real too: Recursion's public RxRx3 release alone covers 17,063 genes profiled across 2.2 million images and 1,674 compounds [5]a, and the company is explicit about the ambition — it describes its datasets as maps of biology, built so that the connectedness of human biology can be navigated toward new medicines [5]a. That is not marketing fluff bolted onto a wet lab; it is a coherent scientific bet that if you generate enough consistent, high-dimensional phenotypic data you can learn relationships that no hypothesis-driven programme would have proposed.
Generative chemistry has become routine rather than exotic, and the paper concedes the achievement while bounding it: published, experimentally validated workflows have raised hit rates and found novel chemotypes, but the successes are, in the authors' assessment, still largely successes of ligand design rather than of drug design [1]a — which is a real capability, correctly labelled.
The clinical numbers, read charitably, are not embarrassing either. AI-designed molecules have reached the clinic, and a survey of the field concludes that they are entering trials at an increasing rate while progressing through development at about the same rate as conventionally discovered compounds [6]a — which is to say: no worse than the industry baseline, from a standing start, in about a decade.
And the counting argument cuts both ways. Bender et al. note that most projects at "AI-first" companies are still preclinical, a few dozen are in phase I or II, and very few have reached phase III [1]a. But they also note the confound in the optimists' favourite statistic — reported phase I success rates of up to 90% — namely that many of those projects are built on well-established biology and chemistry, which lowers the odds of running into safety problems the second time around [1]a. Both readings are unresolved by the data, and the paper says so. Its framing, which I think is exactly right, is that the field is at a stage of "absence of evidence" and not necessarily "evidence of absence" when it comes to AI's translation into clinical impact [1]a.
That is the position this review starts from: not "AI in drug discovery has failed", but "the experiment that would tell us has not been run properly yet, and here is why that is harder than it looks."
2. The fifteen rules
Rule 1 — Don't dismiss AI in drug discovery as all hype. Don't believe everything you read, either.
The authors convincingly argue that the jury is still out. AI obviously has great potential and has been an integral tool far before LLM's came into popular use. However, drug discovery is a long chain of decisions (target choice, assay choice, screening-deck luck, chemist intuition, animal-model availability, etc.), so it's very difficult to isolate why any one project succeeded or failed. This feeds a common bias — failure gets attributed to "bad luck," success gets attributed to "the method" (computational or experimental) — which distorts public and internal perception of what AI actually contributed [1]a.
Rule 2 — Think hard about which stage you are applying AI to. Hint: phase II.
The authors argue that any AI methods that can contribute to patient stratification and ensuring efficacy and safety during Phase II will have the largest impact because they address the longest and most expensive part of the process [1]a.
For example, using biomarkers to stratify patients to the right cohort roughly halves the capitalized cost per successful drug launch (based on ~21,000 compounds in commercial databases, 2000–2015), purely because it raises clinical-phase success rates [1]a. Genetic support for a drug target is cited as another example of an evidence type that measurably raises clinical success odds (drug targets are reported to be ~2.6× more likely to succeed clinically when supported by genetic evidence) [1]a, a figure that traces to Minikel and colleagues [7]a. See Figure 1 in the paper.
On the other hand, much more attention is paid to the use of AI in the preceding Discovery Phase - partially due to the presence of large, homogenous and well-curated data sets that can support this activity [1]a. But the authors point out that the results of these early phase studies (targets and ligands) are not the crux of the problem. There are reportedly >10⁶ known bioactive ligands (ChEMBL, PubChem) versus only ~10³ marketed drugs (see Box 1 in the paper) [1]ab.
Factors where AI can have the largest impact on decision making in drug development include the right drug (physicochemical properties, on-/off-target activity, metabolism and dosage), given to the right patient (patient endotyping and patient stratification) — see Table 1 in the paper.
The authors warn that many choices are locked in early (hit-to-lead), but their consequences aren't visible until years later in the clinic — a fundamental structural challenge for validating any AI model against the outcome that actually matters [1]a.
Rule 3 — Modulate your expectations of what AI can do with biological data.
Biological data (cellular/in vivo disease models, pharmacology, clinical trial data) differs fundamentally from the data types where AI has succeeded. The proverbial image classifier that determines "is this an image of a cat" is "easy" to train because it is possible to easily collect thousands of images from the internet that are objectively "of cats" and thousands that are "not of cats". This pattern does not hold for most biological datasets which are:
highly conditional - outcomes depend on dozens to thousands of variables (genetics, dose, comorbidities, diet, gut microbiome, smoking status via CYP enzyme induction, etc.), so results from a "single" perturbation are rarely reproducible across contexts [1]a and
epistemologically opaque - we frequently lack solid mechanistic/theoretical understanding connecting a given input to a given output, unlike, say, physics [1]a. As a result, it's often impossible to assign clean, unconditional labels to compound effects and this undermines the reliability of any supervised model trained on those labels [1]a. Table 2 and Figure 2 in the paper give examples of these labelling challenges and their consequences.
One response of the last decade has been to manufacture large, homogeneous datasets on purpose, using high-dimensional assays such as Cell Painting or single-cell RNA-seq (Box A). Whether that works is Rule 9 and, separately, the first point under Rule 7.
Box A — Cell Painting and high-content phenomics. Cell Painting is a multiplexed fluorescent staining protocol that labels several organelles at once, so that a single microscopy well yields hundreds to thousands of morphological features per cell rather than one number. Applied across compound or CRISPR perturbations it produces a "phenotypic profile" that can be compared between perturbations without a prior hypothesis about mechanism. Bender et al.'s assessment of where it stands is cool: the recent advances have not yet been validated clinically, and getting models trained on biologically relevant data will take deliberate data generation and quality control [1]a. Refs [1], [5].
Rule 4 — Modulate your expectations of what AI can do with chemical data, too.
Chemical space is enormous (up to ~10⁶⁰ small molecules) and high-dimensional; any real dataset samples only a tiny, biased slice of it [1]a. Chemical space is also local: appending the same functional group to two different scaffolds can produce very different effects, making both sampling and extrapolation fundamentally hard. Figure 3a in the paper uses "chemical space as the universe" as an analogy — a new project's location in that space is, by definition, unknown at the time a model is built, so test-set performance doesn't reliably extrapolate to future projects. The paper flags that "external" test sets in the literature are often not truly external — they're frequently just another subset of the same original data pool — which systematically overestimates real-world model performance [1]a. Figure 3b in the paper shows empirically that four commonly used ADME datasets (solubility, permeability/BBB, clearance) share almost no compounds or scaffolds with each other (overlap as low as 0.1–0.7%) [1]a, meaning each model trained on them has a different, narrow applicability domain — which complicates multi-objective optimization, since drug design always requires balancing many properties simultaneously (see Box 1 in the paper) [1]a. Information leakage from same-source train/test splits artificially inflates reported performance; the paper notes a psychological dimension too — scientists have limited incentive to adopt harder, more realistic (e.g., time-series) splits that make their own results look worse [1]a.
Rule 5 — Not every problem is solvable by AI, even when it has a benchmark.
Benchmarks drove genuine progress in image and speech recognition because performance gains there transferred to real-world use. In drug discovery, the authors argue this analogy breaks down: the conditional, ambiguous nature of biological/chemical data means many published benchmarks are built on an underspecified problem setting [1]a, creating a gap between "we improved the benchmark score" and "we improved real-world decisions." They explicitly name this risk "SOTA-chasing" and invoke Goodhart's law (once a measure becomes a target, it stops being a good measure) — benchmarks can and do get gamed [1]a.
Instead, the authors recommend following benchmarking studies that aren't primarily trying to "sell" a new method are more credible when they use a genuinely relevant endpoint and aren't optimized for glamour. Blind, prospective, community-run challenges — CASP (protein structure), SAMPL (protein-ligand modeling), CACHE (computational hit-finding) — are cited as among the few mechanisms that assess performance rigorously [1]a, though even these usually sidestep the deeper label-conditionality problem and require heavy community coordination plus genuinely prospective ground truth, which isn't always available.
The authors also express the important but often forgotten maxim that generating data is not the bottleneck in drug discovery — "reduction to practice" is (i.e., turning data into clinically relevant decisions) [1]a. The Human Genome Project is commonly used as an example: it gave the field an invaluable static "parts list" of biology, but a 2001 prediction that target identification/validation would become a "high-throughput process" as a result has still not been fulfilled roughly two decades on [1]a.
This idea of the disconnect between static knowledge and the ability to change the world around us is best expressed in one of my favorite passages from the paper [1]ab:
Compared with geography, where the map of the world is largely static and so can be used practically to move from A to B, biological data are heterogeneous between individuals, within an individual and over time, and it also changes as a function of many factors, such as ageing and nutrition. Furthermore, biological data are conditional as noted above. Hence the analogy of building a ‘map’ of biology funda- mentally does not hold. Furthermore, we do not even really know what ‘A’ and ‘B’ are in many cases — for example, when classifying disease or cellular states, it is often not unambiguously defined which markers (such as surface markers) one needs to look for when assigning disease and cellular states to a system. Applying AI tools in the area of protein structure prediction — where we have labelled data — has been fruitful because we have a ground truth. For example, AlphaFold falls into this area, as it is built on about 50 years of data deposited in the Protein data Bank (PDB), where electron density (with all its shortcomings) represents a suit- able ground truth for model generation and evaluation. Such a scenario is not the case for many other desired predictions. — Bender et al. 2026 [1]
Having been one of those researchers that was around in the early 2000's when we were ostensibly "building maps" of how the cell works, this resonates deeply. By all means, encyclopedic biological data in all its forms are critical to the study of life, but trying to translate this piecemeal data into biological consequence is fraught with difficulty because it is context (molecular, cellular, individual) that has the final say in biological function.
Using AI to develop drugs is something to be contrasted (and not confused) with recent advanced uses of AI such as AlphaFold. Where AlphaFold differs is that it had a genuine ground truth to train and evaluate against — as noted above. All local context required to solve the folding problem were implicit in the data available. Most desired drug-discovery predictions don't have an equivalent ground-truth resource [1]a.
Rule 6 — Don't just validate the model; validate its use inside a process.
A model is only useful in terms of the decision process that it facilitates. Most model validation ends after covering the correlation between input data and prediction. The authors argue that real-world value requires two additional links the field often ignores: (1) the project context feeding into the model (disease, endotype, target, anticipated human dose, etc.), and (2) the decision/follow-up experiments the prediction is actually meant to inform [1]a.
The authors provide instructive examples. For project context, a toxicity model trained on one concentration cutoff may simply not be relevant if the concentration ultimately reached in humans differs from that cutoff in a later phase of development [1]a. For decision context, they provide the example of two models with near-identical AUC (a generic, frequently reported metric) that behave completely differently depending on the use case — model 2 is roughly 2× better in an early selection/hit-finding setting, while model 1 is roughly 3× better in a late-stage deselection/toxicity-screening setting [1]a. The takeaway: generic metrics like AUC can be actively misleading, and reporting them without reference to a use case is a real contributor to why published performance and real-world performance diverge.
The paper's recommendation is to move from just model validation to model and process validation — "models need to be linked to processes they are used in" [1]a.
The authors reference a gem (All models are wrong and yours are useless) that poignantly describes why this matters in real life [8]ab. The whole article is worth reading. Here are a few snippets:
Some years ago my medical collaborators and I proposed a new classification of colon cancer into subtypes and helped to consolidate competing classification schemes. To broaden applicability, we even designed histopathology markers for the subtypes. This was a lot of work and we did our best to make the subtypes accessible and useful, but as far as I can tell, none of these classification schemes gets regularly used on patients. And the colon cancer subtypes are not an exception, the subtypes my lab helped to define in breast and pancreatic cancer are also not being widely used clinically. . . .
Observation 1: Success in academia is not the same as success in the clinic. . . . As my own examples above show, academic success does not necessarily lead to clinical success. Why? Because there is little incentive to actually implement an academic advance. Academic career rules prioritise novelty over implementation. . . .
Observation 2: Successful models use data that are available in routine practice. Not just the incentives, also the data differ between academia and the clinic. Large academic collections like TCGA make it look like integrating DNA with RNA with methylation with imaging with proteomics was already general practice - whereas in fact the only data you might have in clinical reality are an H&E slide and some DNA, hopefully from the same patient. . . .
Observation 3: Successful models are linked to actions. The reason the cancer subtyping studies I described above lack the impact I was hoping for is that they indicate differences in survival without being linked to a clear action. Some people do better and others do worse, so what? Similarly, the original PAM50 classifier for breast cancer subtypes had no action linked to it and was useless until the ProSigna test modified it into a prognostic score to recommend adjuvant chemotherapy to high-risk patients. What doctors really need to know is: What action should we take to help a particular patient? What drug, if any, should we give them? These are the questions your prediction tool needs to address to even have a chance of being useful. And the best way to find out if you are addressing an important clinical decision point is to engage closely with a wide variety of clinical practitioners and domain experts.
-- Markowetz, F. All models are wrong and yours are useless: making clinical prediction models impactful for patients. npj Precis. Oncol. 8, 54 (2024) https://pmc.ncbi.nlm.nih.gov/articles/PMC10901807/ [8]
Rule 7 — Treat "AI in drug development" as a stack of separate, unevenly mature fields.
AI in popular culture is commonly treated as a single "thing" while in reality it is a collection of methods applied to a collection of problems. It's only useful to talk about AI in drug development (like "computers" in drug development or "math" in drug development) in as much as one can identify things common to its practice and development across these subfields. This paper accomplishes that.
The "Where are we?" section (Table 3 in the paper) provides a brief overview of these subfields, dividing them between the molecules they are applied to (small molecules, proteins, peptides, antibodies), application to treatments (cell and gene therapy), data types (text and image data), functional assay development (cell-painting assays, temporal/single-cell/spatial transcriptomics) and translational data (real-world evidence, LLMs for diagnosis, trial-design, patient recruitment and response prediction).
The coverage is obviously non-exhaustive, and the authors say so themselves: the state of the art inside companies may well be further along than the published record, and they decline to judge what has not been published or clinically demonstrated [1]a. What's important is that each subfield has its own distinctive challenges, limitations, state of maturity and recent advances. It's a non-homogenous mix. The one diagnosis that cuts across all of it is Rule 6's: the authors want more attention paid to how data, models and use cases are joined together [1]a.
A few points were of specific interest to me. First, the authors consider that High-content imaging (Cell Painting, microscopy) used to train functional models for phenotypic drug/target discovery and disease modeling hasn't yet reached clinical validation, and needs strong data-quality control and biological diversity in training data to generalize [1]a. The authors seem to hold the same view of use of temporal, single-cell, and spatial transcriptomics. These techniques enable finer-resolution characterization of biological systems and compound effects than before but the open challenge is going from high-dimensional data to variables that are actually useful for clinical decision-making [1]a. There are companies that champion these exact methods — Recursion, whose public RxRx3 phenomics map is built from image-based, high-dimensional data generated in its own automated labs [5]a; insitro, whose CellPaint-POSH platform (POSH: pooled optical screening in human cells) combines pooled CRISPR screening, high-content imaging and machine learning to map gene function without a prior hypothesis [9]a; Relation Therapeutics, whose "lab-in-the-loop" combines tissue profiling, single-cell and spatial transcriptomics with target validation [10]ab; and Noetik, which trains foundation models on spatial transcriptomics of tumours and offers to infer spatial transcriptomes from ordinary H&E slides [11]a. I think there is an argument that they take a disproportionate share of the attention space around AI and drug development and bear some responsibility for the popular impression that AI is just about finding drug targets - something this paper attempts to correct. A detailed review of these ongoing efforts is beyond the scope of this review but something I think should be addressed in a future post. For the time being, I would concur with the authors' carefully guarded statement that we are at a stage of "absence of evidence" and not necessarily "evidence of absence" [1]a.
Secondly, the authors are explicit that LLMs "solely process patterns in language on a statistical basis, rather than 'reason'" [1]a. This is true of LLM's but not necessarily true of AI systems that are agentic neural-symbolic systems - I've written about this before in my blog on agentic programming in practice [12]a, [13]a, [14]ab. Notably, the paper explicitly flags agentic LLMs for drug-discovery tasks — while acknowledging they've been described in the literature, the authors caution that scoring well on such agentic tasks may not translate into better real drug discovery outcomes, because the underlying tasks are often underspecified and don't represent real-world project conditions [1]ab. And this is really all the authors have to say about this topic which seems odd (and almost dismissive) given that it is recently a very active sub-field of AI and drug development and that Bender himself is a co-author on a recent review of this area [3]a — one which, incidentally, treats the "reasoning" capacity of the models at the heart of agentic systems as an open question rather than a settled negative [3]a - I can only speculate that the authors decided further discussing the topic would be distracting from their present argument.
Rule 8 — Decide what to measure before you decide what to model.
With Rule 7 the review of "where we are" ends and the paper's own prescriptions begin. The authors frame them as two phases, in order. The first job is to find out where today's data fall furthest short of predicting what happens in clinical development [1]a — and only then, second, to refine how models are built and tested against the use they will actually be put to, productionise them, and scale. The ordering is the whole point: the authors would rather the field did the right thing imperfectly than the wrong thing with technical polish [1]a, which they underwrite with Tukey's line about preferring "an approximate answer to the right question" to "an exact answer to the wrong question" [1]a.
They compress the target into a value equation — Model value = improvement in clinical success rate × project applicability domain — whose second factor is the underappreciated one: a model is worth more the more projects it can be applied to [1]a. A generic safety or PK model that applies across a portfolio is worth more than a brilliant target-specific one, all else equal. Which is an uncomfortable finding for a field whose publication incentives reward the brilliant specific model.
The failure mode this rule guards against is named too: datasets generated because a technology made them possible ("tech push") and only later repurposed for a question, as opposed to datasets generated because a scientific question demanded them ("science pull") [1]a.
Rule 9 — Generate data in systems that resemble the human you are trying to treat.
If the bottleneck is data predictive of human outcomes, then the recommendations are about what to run, not what to fit. The paper's list, in ascending order of ambition:
Build in vitro and ex vivo assays with high-dimensional functional readouts — Cell Painting (Box A), single-cell RNA-seq — and use ML to discriminate genuinely different physiological cell states, rather than collapsing to one low-dimensional number. The worked example is classifying inflammasome-perturbed cellular states instead of measuring the single cytokine readout IL-1β alone [1]a.
Use human primary cells, ideally from many donors, rather than immortalised lines or animals, so that preclinical testing captures more of the variability between people and finds drugs that work across a broad population [1]a.
Take cell-type-specific vulnerability seriously — motor neurons in ALS, dopaminergic neurons in Parkinson's, non-cell-autonomous microglial and astrocytic effects in Alzheimer's — and prefer human iPSC-derived models where relevant. This is not aspirational: the paper's example is iPSC-derived motor neurons from ALS patients, used to find disease phenotypes and to test ropinirole and bosutinib in vitro, a line of work that has already produced clinical candidates [1]a.
Treat virtual cells as a long-term goal — computational systems that might one day link a disease or drug perturbation to a phenotype-level outcome [1]a.
Longer term still, a causal genetics framework: large-scale genome editing in human cell systems chosen to match the disease, so that perturbation, context and phenotype line up and disease hypotheses can be tested directly, cause to effect [1]a — extended beyond monocultures into co-cultures that capture crosstalk, such as cancer cells with their microenvironment.
The overarching instruction is the one that indicts a fair amount of the last fifteen years of omics: piling up ever-larger datasets is no longer sufficient; what is needed is data that expose mechanism and bear on clinical outcomes [1]a.
Note the tension with Rule 3 and with the first point under Rule 7, because a reader would raise it right here: Rule 3 says high-dimensional phenotypic data has not yet delivered clinical validation, and Rule 9 recommends generating more of it. That is not a contradiction, but it is only consistent under one reading — that the problem with phenomics to date is experimental design and use-case linkage, not dimensionality. Whether that reading is right is, as far as I can tell, the single most important open empirical question in the paper, and it is not settled by anything in it.
Rule 10 — Close the loop between preclinical and clinical data, including the failures.
Pharmaceutical companies routinely separate preclinical and clinical functions — different databases, different access rules, different people. That separation is fatal to the one thing that would tell you which preclinical assays are actually predictive, because you only find out from clinical outcomes. The paper recommends explicit iterative feedback: phase I data to train tolerability, toxicity and PK models; phase IIb proof-of-concept data to refine target validation and efficacy prediction.
The part I would underline is the instruction to feed the loop with negative results, and the second-order benefit claimed for it: models trained on the failures and the marginal outcomes as well as the successes should translate better, and might also blunt the confirmation bias that steers clinical pipelines [1]a. Not just better models — a partial antidote to the confirmation bias of an entire pipeline. That is a bigger claim than it looks, and it is the one recommendation in the paper that costs nothing technically and everything politically.
Rule 11 — Validate at the decision point, in the workflow, with the metric that aids a meaningful decision.
This is Rule 6 turned into an operating instruction. Current validation is often model-centric, with metrics untethered from use; the fix is to validate models against their ability to aid specific project decision points [1]a — evaluated inside the workflows they are meant to support, with realistic data availability and update cycles, and metrics chosen to reflect the actual decision and its trade-offs, which differ between early discovery and late development. It also means being honest about which experimental follow-ups are actually feasible and whether the provided data will be actionable.
Embedding models in DMTA cycles is the direction of travel, and the paper names names — Iktos, XtalPi, Insilico Medicine, Genentech and Tempus among the companies running AI-integrated DMTA loops [1]a, some of them explicitly incorporating patient data and organoid models. The requirement that makes or breaks it: the selection criteria inside the loop have to predict the human situation, so that automating the loop improves the quality of decisions and not merely their number [1]a. Automation that makes bad decisions faster is not an improvement.
Rule 12 — Solve it at the scale of the problem: consortia, or a purpose-built organisation.
No single company can cover the ground: the chemical space, the number of proxy endpoints and assays, and the number of in-vivo-relevant endpoints those assays would have to be tied to are together beyond any one organisation [1]a. So the first path is industry-spanning consortia, reaching into neighbouring sectors — agrochemicals, consumer products — that work the same chemistry.
The paper is usefully specific about why the previous generation fell short: earlier consortia treated how the data were designed and generated, and the context they came from, as an afterthought [1]a. Data got pooled retrospectively — clearance values combined across companies despite different assay formats, species and thresholds — silently assuming an equivalence that did not exist. The fix is that companies must proactively define needs and design experiments for harmonised training, supported by cross-industry and government funding.
The second path is a dedicated organisation that generates purpose-built data after a first, exploratory phase has settled what the questions are. Three real examples are given: the OpenADMET consortium, funded by ARPA-H at US$30 million to generate data on off-targets linked to clinical liabilities [1]a; and LIGAND-AI and Ginkgo Bioworks' Virtual Cell Pharmacology Initiative, covering ligand–target interaction data and compound DRUG-seq readouts respectively [1]a — DRUG-seq being the high-throughput transcriptomic profiling of compound perturbations introduced in §1.2. Two publicly funded, one industrial — which is the point: both routes work.
The summary imperative is the paper's thesis in one line: the field has to keep shifting its attention from the technology to the science of what actually needs doing [1]a — with more weight on approaches that are process-validated, that scale, and that experimental scientists actually use, which requires attention to interface design, education and organisational culture rather than modelling technique alone.
Rule 13 — Account for the incentives: yours, your company's, and your journal's.
This is the section you don't normally get in a review journal, and it is the reason I would recommend the paper to people who have no interest in cheminformatics.
The macro pressure is Eroom's law (Box B): the ever-rising cost of getting a drug approved, which the authors name as the contextual force pushing pharmaceutical companies to change how they work [1]a. Against that backdrop, "AI-first" companies tend to launch on extravagant promises, and the pressure on company and investors alike to deliver on those promises can distort the decisions that follow [1]a. The authors also suggest, mildly, that the emotional biases of decision-makers inside pharma are underappreciated as a factor [1]a. Anyone that has worked in biotech will know that adopting a new technology (even if it's just a small part of larger pipeline) comes with significant cost that needs to be weighed against the potential long-term benefits.
Academic publishing gets its own indictment, and it lands closer to home: models that nudge the state of the art on toy datasets get published and noticed, while their bearing on any real process may be negligible or simply never examined [1]a. Combined with insufficient attempts at real validation, the authors argue, this manufactures a sense of progress that overstates what has reached practice. The example cited under Rule 6 resonates with this point.
Box B — Eroom's law. Moore's law spelled backwards: the observation that the inflation-adjusted cost of developing one approved drug has risen roughly exponentially for several decades, even as the underlying science and tooling improved. It is the economic backdrop against which every efficiency claim in this field has to be read, because it means the industry has a strong prior incentive to believe a new technology will work. Bender et al. invoke it as the contextual pressure behind the industry's appetite for process change [1]a. The original analysis is Scannell and colleagues (2012), linked from the glossary entry in the Appendices. Refs [1].
Rule 14 — Respect emergence: reductionism plus more data is not a plan.
Most biological research proceeds by disassembling a system into components and assuming that a component tuned in isolation will behave the same way once it is back inside the whole. The paper's verdict is that the assumption routinely fails — redundant pathways, feedback loops and cell–cell interactions produce emergent behaviour that the parts list did not predict, and the authors count this among the drivers of clinical failure for lack of efficacy [1]a. The canonical example is target inhibition that clearly modulates a pathway in vitro or in an animal, and is blunted by compensation once it meets a human.
Multiomics and systems biology emerged partly in response. But the paper argues that the hard problem simply relocates: to choosing which variables to measure, in which biological system, and at what point in the course of disease and treatment [1]a. And even with the data in hand, extracting actionable conclusions is limited by the fact that measured variables routinely outnumber informative samples. The authors suggest that the solution will somehow involve not more data but more structure: folding prior biological knowledge into the inference — through Bayesian or causal models, or through Mendelian randomization — is, in their view, probably unavoidable [1]a.
Context is everything. This bugbear of biology underlies almost every rule above and may prove the reason that whilst AI will prove a useful tool, its perceived promise to "solve" all diseases is an illusion.
I would add one last question. If LLMs can be made to appear to reason more effectively over text by constraining that output with deterministic code that is relevant to the immediate context of a conversation, is there an equivalent in biology? If reductionist thinking over biological facts is akin to the LLMs working over text, what is it that can be added to this that is akin to constraining, deterministic code that is relevant to the immediate context of a biological line of thought?
Rule 15 — Pick numbers that matter.
The authors' closing demand is that the field measure itself by "numbers that matter" rather than by proxies contrived to generate publications instead of R&D progress [1]a. And the numbers that matter are named — higher clinical approval rates, and genuinely new therapeutic modalities and mechanisms of action [1]a — not modelling novelty.
The authors frame the Perspective's goal modestly, as contributing a realistic view of where AI in drug discovery currently stands so as to clarify the conditions under which it can deliver. On the evidence of the fifteen rules above, they have done that.
3. How this article was written
My primary interest in this article was as background for a larger project to survey design patterns in Agentic AI systems used for drug development and biological research. My secondary aim was to explore the use of an adversarial agentic AI system I had created to debate positions based on published literature. The system is composed of a writer and critic agent that start with a human written position and take turns to revise the position in subsequent back-and-forth rounds until some quality metric is achieved.
I wrote the starting position based on my own reading of the paper and an attempt to distill its point into a set of rules. After a single round, I expanded on the paper's treatment of LLMs in drug development (the coverage is light here) to include my own questions and ideas around use of neuro-symbolic systems in drug development as well as a second paper with Bender as co-author that covers AI agents in drug discovery in far deeper detail (see Huynh, Seal and colleagues' AI agents in drug discovery: applications and case studies [3]). The final machine version was preceded and followed by a human review. That expansion outgrew this article and has been set aside as the seed of a follow-on post on generative and agentic AI in drug development.
Quotations from these two papers were woven into the first drafts of the piece and were checked by the critic agent. Because a justified two-column PDF's text layer carries soft hyphens, superscript citation numerals and kerning artefacts that are not part of the authors' wording, the critic agent used a deterministic script for quote checks. Across those drafts it reported 105 verified, 0 failed across the two PDFs.
For this version the citation style changed. The prose now paraphrases, each claim taken from a source carries a numbered citation, and where a reader might want to cross-examine a claim the citation is followed by a provenance footnote that links to the exact passage on the publisher's open-access page (Bender et al. is now open access at nature.com). A script fetches every footnoted page, reduces it to plain text and checks that the footnoted passage is a verbatim substring of it — the current result, 88 verified, 0 failed (86 web fragments and 2 page anchors into PDFs whose publisher pages refuse scripted access), is reproduced in the Provenance and verification section of the Appendices. A second script checks that, outside marked quotations, the article shares no run of eight or more consecutive words with the paper it reviews; its output is there too. Validation data was collected to report files for human review. The provenance footnotes are debate scaffolding and will be removed before publication.
Every number in the prose carries a footnote naming where it came from.
References, glossary and verification record
The numbered references, the glossary of terms used above, and the provenance and verification record for every citation are in a companion post: Appendices.
