The distinction this paper is about

A benchmark measures whether a stage is right on the data it was tested on. It does not measure whether, when that stage is wrong on new data, anything in its output reveals it. Those are different properties, and the second decides whether a pipeline can be trusted without ground truth, which is the situation genome mining is always in.

A method that returns "unknown" for a protein it cannot handle is behaving safely: the failure is legible and a downstream step can route around it. A method that returns a specific, confidently scored, wrong function is behaving unsafely, because its output is formally identical to a correct one.

What we did

We audited a working natural-product genome-mining system of our own. We did not design experiments to elicit these failures; all seven surfaced from a single discipline of re-deriving every number in a manuscript from the files that produced it, rather than from the summary that reported it.

What we found

An annotation metric that counts vocabulary, a union and intersection at chance, and cross-tool overlap

The keyword rule scores one genome 0.0% while 1,712 of its 3,167 products are blank. Two tools' recoveries union to 356 of 2,455 proteins (14.5%) and intersect at 11 against 10.38 expected under independence (Fisher p = 0.87). Across 58 genomes only fungi have the coverage to test agreement.

Seven stages where wrong output was indistinguishable from correct. Five were detectable only by recomputation. Three patterns recur:

  • Metrics measure something adjacent to what they claim. An annotation-quality rule counting "hypothetical protein" scored one genome 0.0% where 54.1% of its products were empty: the genome is annotated in a different dialect, and the keyword rule measured which words a submitter chose for their unknowns.
  • Evaluations cannot save their own confound. A classifier scoring one protein 0.988 then 0.000 at unchanged ROC-AUC 0.995 reproduced neither label value once rebuilt with documented classes and seeds. The same confound cut a published classification accuracy from 70.4% to 49.5% on unseen genera.
  • Provenance records describe the invocation, not the experiment. A seven-method benchmark spanning BLAST/Pfam, Foldseek, CLEAN, ESM-2, ProtRek and DeepFRI graded all seven wrong against a ground truth taken from an unverified code comment. Four were in fact correct. The protein had been misidentified, and correcting it inverts the reading.

A combined recovery figure reported as 15-20% was 14.5% over the 2,455 proteins both tools actually scored, and the intersection the figure implied had never been measured: 11 proteins, a Jaccard index of 0.03.

Seven methods on MA_4560 regraded, and a classifier whose score was not about the protein

Figure 2. A benchmark whose key was wrong. (a) Seven methods on MA_4560, each cell the prediction actually returned. Graded against a ground truth taken from an internal code comment, all seven read as failures; graded against UniProt and InterPro, four return the assignment those databases support and three correctly abstain. The predictions never changed, only the key.

Embedding cosine against compound Tanimoto, and the same embeddings under retrieval

Figure 3. A relationship no single coefficient describes. (a) Embedding cosine against compound Tanimoto over 79,401 MIBiG pairs: Pearson r = 0.037. (b) The same embeddings under nearest-neighbour retrieval: mean Tanimoto 0.278 against 0.124 for random pairs, a 2.24-fold lift. Both results are correct and support opposite summaries.

Where each of the seven failures sits in the pipeline and whether its own output reveals it

Figure 4. Where each failure sits, and whether the stage's own output reveals it. Five of the seven produce output that is internally consistent and indistinguishable from success. Each was found by re-deriving a number from its inputs, not by inspecting a result.

What this does and does not show

Three claims from our own prior work are withdrawn in full here, including a reported instance of fold-convergent annotation-transfer failure that we no longer have. The paper reports failures in a snapshot of actively maintained tools, so version numbers are not the subject and no claim is made that archaeal biology causes any of them. We release the checks so the audit can be repeated elsewhere.

PDF · Supplementary