The question
Computational natural-product discovery is organized around the biosynthetic gene cluster, an arrangement characteristic of bacteria. MIBiG v3.0 holds 3 archaeal entries among 2,502. Whether archaeal biosynthetic capacity is genuinely absent, under-detected because archaeal genomes are organized and annotated differently, or limited by genome mining generally, had not been assessed against matched bacterial controls.
What we built
Detection status for 12 archaeal compound classes; 3,935 classified BGCs across 1,681 genomes (terpene 42.8%, unknown including TfuA 21.4%, antibiotic-type 16.9%); distribution across eight phyla, unnormalised for sequencing effort; and archaeal against bacterial thiopeptide identity at 26-39%, below the ~40% homology-tool threshold.
- A survey. 3,974 BGCs across 1,681 archaeal genomes and eight phyla. Of the 3,935 assignable to a functional category: terpene clusters 42.8%, clusters of unknown function including TfuA-associated loci 21.4%, antibiotic-type clusters 16.9%, all other categories 18.8%.
- BGCast, a contrastive-learning substrate predictor for adenylation domains that aligns binding-pocket embeddings with substrate fingerprints. Trained on 8,744 bacterial domains, it reaches 74.7% ± 0.3 top-1 and 84.1% ± 0.4 top-3 across 54 substrate classes, against approximately 85% for antiSMASH.
What the controls showed
Two candidate explanations for the detection shortfall were tested against matched bacterial comparisons, and neither survived:
- That the pathways are not organized as detectable clusters. The bacterioruberin pathway of Haloferax volcanii spans 81% of its main chromosome, but the same keyword signature recovers comparable dispersion in three bacterial carotenoid producers whose pathways are genuinely clustered.
- That the genes lack functional annotation. Archaeal proteomes average 40.2% hypothetical, against 32.1% across eight bacterial genomes spanning curation intensity. Curation intensity predicts the fraction better than domain does.
The shortfall therefore locates in general properties of genome-mining methods and in curation depth, not in archaeal genome biology.
Figure 2. The two usual explanations, tested. (a) Bacterioruberin biosynthesis in H. volcanii spans 17 loci across 2,298,501 bp, 81% of the main chromosome, but only 4 are carotenoid-dedicated, and matched bacterial carotenoid producers show comparable dispersion. (b) Archaeal proteomes average 40.2% hypothetical against a matched bacterial control at 32.1%, which does not significantly differ.
Figure 3. A training-population artefact. MA_4560 is not TfuA; it is an uncharacterized FilR1-family regulator. The classifier scored it 0.988 against archaeal housekeeping negatives and exactly 0.000 once bacterial negatives were added, at unchanged ROC-AUC: it had learned bacterial-versus-archaeal, not biosynthetic-versus-not.
Figure 4. BGCast. (a) Frozen ESM-2 read at 34 binding-pocket positions, aligned to chirality-aware Morgan fingerprints by symmetric InfoNCE. (b) 74.7% top-1 and 84.1% top-3 over 54 substrate classes on a held-out split, against roughly 85% for antiSMASH. These are upper bounds: train and test are not disjoint at the level of the model's input.
Figure 5. The roadmap, with its arithmetic on the panel. Of 12 literature-curated archaeal compound classes, 1 is fully detectable, 3 partially, 2 have no rule in any tool and 4 have a rule of untested archaeal fit, giving 20.8 to 54% depending on how that last group is credited. Deliberately a range, not a single number.
What this does and does not show
The 74.7%/84.1% figures are upper bounds: train and test splits are not disjoint at the level of the model's input, and 74.0% of test domains carry a 34-residue pocket string byte-identical to a training domain. BGCast is architecturally capable of ranking substrates absent from its training data, but the class-count filter removes every unseen class from the test split, so that capability is not demonstrated here. Detection rules cover an estimated 21-54% of 12 literature-curated archaeal compound classes. The survey is weighted toward halophilic and marine lineages rather than sampling archaeal diversity evenly, so the genome counts are a lower bound over an unrecorded screened population, not a census.




