Premium accounts now available! Sign up and create a premium account. Read more Close

Advertisement

Image

The phage stress test: a proving ground for genome language models in biological prediction

Preprint Created on 24 Sep 2026 bioRxiv

Reliable biological prediction by AI is most likely to emerge first in the simplest systems with rich datasets and the fastest opportunities for testing and refinement. Phages with small genomes are an ideal testbed. Decades of experimental work have created an unusually information-rich literature benchmark that sits largely outside the sequence repositories typically used to train genome language models (GLMs). Recent work has shown that GLMs such as Evo2 can generate viable whole bacteriophage genomes, demonstrating that genome-scale biological design is possible. The next question is more mechanistic: can such models correctly predict the effects of simple, local sequence changes? Here, we evaluated Evo2 against experimental mutation data from two model phages: the single-stranded RNA phage MS2 and the single-stranded DNA phage {Phi}X174. We first addressed whether Evo2 sequence log-likelihood scores could distinguish viable from nonviable mutations, and then whether models trained on Evo2 embeddings were predictive of function. Across both phages, Evo2 captured expected sequence-level constraints: stop codons were generally penalized, synonymous substitutions had higher likelihood than nonsynonymous substitutions. Evo2 distinguished among synonymous codons in ways only weakly explained by host codon usage. However, for MS2, these capabilities did not translate into robust biological predictions. Specifically, Evo2 {Delta}SLL failed to distinguish functional from nonfunctional mutations in the lysis gene, an overlapping viral region that appears to be a particularly hard test case. In {Phi}X174, performance was stronger but still modest, and much of the apparent signal could be explained by simple covariates such as nonsense mutations, genomic position, and nucleotide distance from the reference. Together, these results introduce phages as a tractable proving ground for stress-testing GLMs against experimentally grounded genotype-to-phenotype tasks. More broadly, they provide a durable framework for identifying what data and model ingredients are required for biologically reliable prediction.

Layton, E. M., Bernauer, M. L., Weinstock, L. D., Small, E. M., Geiselman, G. M., Bachand, G., Cahill, J.

Advertisement

Stats

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 6
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement