The question
Computational natural-product discovery is organized around the biosynthetic gene cluster, an arrangement characteristic of bacteria. MIBiG v3.0 holds 3 archaeal entries among 2,502. Whether archaeal biosynthetic capacity is genuinely absent, under-detected because archaeal genomes are organized and annotated differently, or limited by genome mining generally, had not been assessed against matched bacterial controls.
What we built
- A survey. 3,974 BGCs across 1,681 archaeal genomes and eight phyla. Of the 3,935 assignable to a functional category: terpene clusters 42.8%, clusters of unknown function including TfuA-associated loci 21.4%, antibiotic-type clusters 16.9%, all other categories 18.8%.
- BGCast, a contrastive-learning substrate predictor for adenylation domains that aligns binding-pocket embeddings with substrate fingerprints. Trained on 8,744 bacterial domains, it reaches 74.7% ± 0.3 top-1 and 84.1% ± 0.4 top-3 across 54 substrate classes, against approximately 85% for antiSMASH.
What the controls showed
Two candidate explanations for the detection shortfall were tested against matched bacterial comparisons, and neither survived:
- That the pathways are not organized as detectable clusters. The bacterioruberin pathway of Haloferax volcanii spans 81% of its main chromosome, but the same keyword signature recovers comparable dispersion in three bacterial carotenoid producers whose pathways are genuinely clustered.
- That the genes lack functional annotation. Archaeal proteomes average 40.2% hypothetical, against 32.1% across eight bacterial genomes spanning curation intensity. Curation intensity predicts the fraction better than domain does.
The shortfall therefore locates in general properties of genome-mining methods and in curation depth, not in archaeal genome biology.
What this does and does not show
The 74.7%/84.1% figures are upper bounds: train and test splits are not disjoint at the level of the model's input, and 74.0% of test domains carry a 34-residue pocket string byte-identical to a training domain. BGCast is architecturally capable of ranking substrates absent from its training data, but the class-count filter removes every unseen class from the test split, so that capability is not demonstrated here. Detection rules cover an estimated 21-54% of 12 literature-curated archaeal compound classes. The survey is weighted toward halophilic and marine lineages rather than sampling archaeal diversity evenly, so the genome counts are a lower bound over an unrecorded screened population, not a census.