The recipe under audit
Short chains are taken from PDB complexes on the assumption that a short chain in a complex is a disordered region. Secondary structure is read from depositor records and mapped to three states. Independence is established by sequence identity, or by deposition date when an independent test set is wanted. A model is credited with using the partner when it outperforms an ablated model trained without one.
Each step is a shortcut with an obvious justification, and none of them had been audited against independent evidence on the same data.
What the audit found
- The selection rule. 22.4% of samples have curated support for being disordered; 41.2% have a record with no curated disorder or MoRF at all.
- The label alphabet. Polyproline II has no three-state code, so it is read as coil. 17.9% of coil residues in proline-rich regions are PPII against 5.9% elsewhere, a 3.1-fold enrichment inside PxxP regions.
- The split. 60% of post-2021 depositions share a super-cluster with an earlier entry, so a date split leaks. Removing them costs 60% of the candidate holdout: this is the one step whose correction has a measured price.
- The ablation gap does not measure partner use. Supplying the partner sequence moves pooled AUC-ROC by +0.0035, which is 0.74 times the standard deviation of retraining the same configuration over 45 runs. A within-model partner swap avoids that variance and exposes the real problem: composition-chosen substitutes are near-cognate, with 2,726 of 4,768 complexes taking their decoy from the true partner's own accession. A wrong partner costs 0.0076 in 15 of 15 paired runs.
Where an ablation gap is offered as evidence of partner use, the paper's position is to report it with its decoys stated, alongside a calibration task with a known partner-dependent rule. Their model recovers 84-90% of a coarse form of that rule; on a finer form, recovery is seed-dependent.
What this does and does not show
The benchmark is evidence about models on this benchmark, not about proteins. For five of the six findings the paper can show the step is wrong without showing how far any downstream conclusion moves; only for the temporal split is the cost of correcting it measured. The sixth is a limit of the data rather than a step that can be corrected.