Read the original source ↗

The problem with the published numbers

Models that predict clinical trial outcomes report AUCs of roughly 0.72 to 0.88. Those numbers come from evaluation splits that let the same compound appear in training and test. A model can reach them by learning which drugs tend to succeed, without learning why a mechanism fits a disease.

This paper names that gap and measures it. Retrained under a fair split that scores only on drugs absent from training, the leading prior model (HINT) falls to 0.626. The highest reported figure in the literature, 0.882 from inClinico, held only under a looser evaluation that included a reverse-causation trial-count feature.

What MERIT does

MERIT trial-selection funnel and model architecture

Figure 1. 428,377 AACT records narrow to the modelling cohort of 3,133 trials across 753 compounds (PASS 2,558, FAIL_EFFICACY 490, FAIL_SAFETY 80, FAIL_BOTH 5). Eight pre-trial modules plus a gated ninth trial-design block combine into 285 features and two task heads.

MERIT predicts from the biology of the drug and the disease, never from compound identity or development history. It links a drug's intended and off-target effects to tissue-specific efficacy and safety through large-scale drug-protein, protein-metabolite and immune interaction maps.

Eight compound-level, pre-trial modules feed it: tissue-specific binding, pathway engagement, target-disease mechanism and genetics, binding specificity, safety pharmacology, pharmacokinetics, and disease and trial context. No single module carries the model; standalone AUCs run 0.682 to 0.707.

What it found

ROC curves, per-fold AUC and signal decomposition

Figure 2. Compound-holdout ROC: overall 0.770, safety 0.784, efficacy 0.765. Per-fold spread across 25 folds. Signal decomposition: flags alone 0.615, plus disease difficulty 0.689, plus molecular mechanism 0.740, full model 0.770.

  • Honest AUROC of 0.770 across 753 small-molecule drugs and 3,133 trials: 0.765 for efficacy, 0.784 for safety.
  • 0.704 against the retrained HINT's 0.626 on the same fair split.
  • Efficacy and safety follow opposite logic. Efficacy is a matching problem between mechanism and disease. Safety is a single-breach problem. This is why one model predicting both needs to treat them separately.
  • Run in reverse, MERIT nominates indications. It recovered the eventual approved indication for 83% of failed drugs, from the mechanism alone.
  • 55 locked, outcome-blind predictions for drug-indication pairs in ongoing Phase III trials, registered before the readouts, establishing a prospective cohort.

Single-feature fingerprints of three confident misses, and disease-pathway overlap by outcome

Figure 3. Why individual predictions miss. (a) Three confident misses as single-feature fingerprints against the cohort: nilotinib's cardiac-target burden (z = +3.2), panobinostat's epigenetic-mechanism fraction (z = +8.2), metoprolol's brain-partition mismatch against COPD (z = +3.3). Each is strong for that one drug and rare across the cohort, so it explains a single miss without raising population AUC. (b) Disease-pathway overlap by outcome.

Three tests of whether the mechanism score ranks the approved indication highest

Figure 4. From failure to rescue. Three tests of whether the mechanism score ranks the working or approved disease highest: within-drug by mechanism-fit 75% (P = 9x10^-13), within-drug on the independent Open Targets genetic axis 82%, known-rescue recovery 83%. Run in reverse, the same model nominates indications.

What this does and does not show

The 0.770 is retrospective and computational. The prospective cohort is registered but has not read out, so nothing here is yet a demonstrated forward prediction. PASS in the trial labels denotes Phase III/IV completion, not approval or demonstrated efficacy. The indication-nomination result recovers indications that were eventually approved; it does not show that following MERIT's nomination at the time would have produced them.

Read the preprint on bioRxiv ↗ · PDF