A funny thing happened on the way from autocomplete to biology: the autocomplete started writing genomes that got tested in the lab. Not protein snippets. Not a cute motif that makes a reviewer say nice things before asking for four more experiments. Whole bacteriophage genomes, generated by genome language models, then checked against reality, the one benchmark suite that still refuses to be gamed by prompt engineering. Science lists the paper as Generative design of bacteriophages with genome language models, and the title is doing a lot of honest work here. The counterintuitive bit is not that AI can read DNA, because sequence models have been chewing through genomes for years like a raccoon in a biotech dumpster. It is that these models are now being positioned as design tools for living systems, which is a much stranger and more consequential claim. ## What Science and bioRxiv say actually changed Science identifies the work as Generative design of bacteriophages with genome language models, while the bioRxiv abstract describes the technical claim more explicitly: genome language models were used to generate viable bacteriophage genomes. According to bioRxiv, the researchers used Evo 1 and Evo 2 to generate whole genome sequences with realistic genetic architectures and desirable host tropism, using the lytic phage ΦX174 as the design template. Experimental testing of the AI generated genomes yielded 16 viable phages, which is the part where the story exits the slide deck and enters the biosafety cabinet. That workflow matters because biological function often lives in interactions across an entire genome, not in one heroic gene wearing a tiny cape. The bioRxiv abstract frames this directly, saying many important biological functions arise from complex interactions encoded by whole genomes. In ML terms, this is moving from classifying the recipe to proposing a cake, baking it, and discovering it does not taste like a PDF. ## Why viable is the load bearing word The bioRxiv abstract reports several validation steps beyond simply recovering phages. Cryo electron microscopy found that one generated phage used an evolutionarily distant DNA packaging protein within its capsid. Multiple generated phages also showed higher fitness than ΦX174 in growth competitions and in lysis kinetics, according to the same abstract. The most application flavored result is the phage cocktail. bioRxiv reports that a cocktail of generated phages rapidly overcame ΦX174 resistance in three E. coli strains. That does not mean phage therapy is suddenly a solved product category, please put the confetti cannon down. It does mean genome scale generation is being tied to functional properties that matter in a real biological design loop. ## The reviewer shaped bucket of cold water PREreview, authored by James Fraser and Joseph Bondy-Denomy, praised the molecular biology and the experimental anchoring, but it also flagged the evaluation problem that ML people should recognize immediately. The review says the manuscript lacks explicit baselines, making it hard to judge whether the success rate of about 16 out of 300 beats lightweight alternatives. In other words, did the giant model navigate sequence space intelligently, or did it mostly find the biotech equivalent of nearby parking? PREreview also notes that the explored sequence space appears to stay close to known natural sequences. That caveat does not erase the result, but it changes how builders should interpret it. Novelty in biology is not a vibes metric, and if the comparison set is weak, a model can look like a wizard while quietly doing nearest neighbor karaoke. ## What builders should take from Semantic Scholar and the paper record Semantic Scholar lists the bioRxiv version as published on 17 September 2025, with Samuel H. King, Claudia L. Driscoll, David B. Li, Daniel Guo, Aditi T. Merchant, Garyk Brixi, Max E. Wilkinson, and Brian L. Hie among the authors. It also summarizes the work as a blueprint for designing diverse synthetic bacteriophages and a foundation for generative design of useful living systems at genome scale. That is a big claim, but the useful lesson is narrower and more practical. For AI builders, the pattern is the product: generate candidates, constrain them with biological priors, test them experimentally, then report not just wins but baselines and failure rates. The glamour object is the genome model, but the durable system is the loop around it. Biology does not care if your loss curve looked elegant; biology is the coworker who replies to your ten page proposal with a single colony plate. The next thing to watch is evaluation. If future genome design papers compare against consensus designs, constrained randomization, ancestral reconstruction, and other simple baselines, we will know whether large genome models are adding genuine design leverage or just expensive autocomplete with a lab coat. Either way, AI biology has crossed an interesting threshold: the model is no longer only reading the book of life, it is submitting edits to the manuscript. ## Sources - Generative design of bacteriophages with genome language models

Sources