Thibaut Dejean, Barbra D Ferrell, William Harrigan, Zachary D Schreiber, Rajan Sawhney, K Eric Wommack, Shawn W Polson, Mahdi Belcaid
5 min
Traditional protein language models (pLMs) are typically trained on individual protein sequences, which limits their ability to capture the complex functional relationships and interdependencies between proteins within a genome. This study addresses whether extending the context window of transformer-based models to encompass entire viral genomes—and incorporating biologically informed sparse attention—can improve the representation of protein function and genomic organization.
The researchers introduced a long-context pLM architecture capable of processing sequences up to 61,000 amino acids. To overcome the quadratic scaling limitations of standard attention mechanisms, they implemented a content-aware sparse attention strategy. This method uses inferred protein-protein interactions (PPIs) as sparsity priors, effectively restricting attention computations to protein pairs likely to interact. The model was trained on the NCBI Virus database using a two-stage approach: first, identifying putative PPIs using an extended-context ESM-2 model, and second, training the long-context model using these interactions to guide the sparse attention mechanism.
The resulting models, LV-3C and LV-5B, significantly outperformed baseline single-protein models (like ESM-2) in perplexity and embedding quality across extended genomic contexts. The models successfully captured long-range dependencies, such as the well-documented evolutionary coupling between DNA polymerase I, ribonucleotide reductase, and helicase. Furthermore, the embeddings generated by these models demonstrated superior taxonomic discrimination and species classification accuracy, suggesting that incorporating genomic context provides a richer, more biologically meaningful representation of viral sequences.
This work bridges the gap between protein-centric modeling and genome-scale analysis. By enabling the analysis of entire viral proteomes, this framework provides a powerful tool for understanding viral evolution, functional annotation, and metagenomic data, where contextual information is often critical for resolving biological function.
BACKGROUND: The transformer architecture in deep learning has revolutionized protein sequence analysis. Recent advancements in protein language models have paved the way for significant progress across various domains, including protein function and structure prediction, multiple sequence alignments, and mutation effect prediction. A protein language model is commonly trained on individual proteins, ignoring the interdependencies between sequences within a genome. However, biological understanding reveals that protein-protein interactions span entire genomic regions, underscoring the limitations of focusing solely on individual proteins. FINDINGS: To address these limitations, we propose a novel approach that extends the context size of transformer models across the entire viral genome. By training on large genomic fragments, our method captures putative long-range dependencies consistent with inter-protein relationships and encodes protein sequences with integrated information from distant proteins within the same genome, offering benefits across downstream tasks. Viruses, with their densely packed genomes, minimal intergenic regions, and protein annotation challenges, are ideal candidates for genome-wide learning. We introduce a long-context protein language model, trained on entire viral genomes, leveraging a biologically informed sparse attention mechanism in which inter-protein links are inferred computationally and used as sparsity priors. Our semi-supervised approach supports long sequences of up to 61,000 amino acids (aa). CONCLUSION: Our evaluations show improved prediction of masked aa and improved downstream discrimination relative to single-protein models and long-context baselines, with additional validation that our inferred links correlate with independently curated interaction resources.
Sam: Which raises generalization. Viral genomes are compact and functionally dense. Would this transfer to eukaryotic genomes, where gene spacing is wider and regulatory elements are scattered?
Alex: That is the limitation the authors point to. They focus on viral genomes because the compactness suits a sparse strategy. Applying it to something like the human genome would likely need a far more complex interaction map, and nothing in the paper shows that it works there.
Sam: Now the embedding space. LV-3C produces a much wider distribution of Euclidean distances than the ESM-2 baseline. Does that dispersion track biology, or is it an artifact of the longer context?
Alex: The evidence points to biology. As taxonomic distance increases, embedding similarity decreases. That is what you would expect if the model were capturing evolutionary divergence rather than noise from a larger window. It also suggests the model integrates genome-specific context, so similar-looking proteins can be separated when they play different roles in different viruses.
Sam: It is still a correlational check, though. And the STRING validation gives an AUC of 0.66. How much weight should that carry?
Alex: Not much for any individual prediction. It is a modest signal, and the authors use it to show the attention is picking up real co-evolutionary structure rather than only training-set patterns. A careful referee would want it paired with a direct comparison against dense attention, which this setup was built to avoid.
Sam: So the load-bearing piece is the architecture: a biologically derived mask that makes whole-genome context tractable. The top-50 analysis and the embedding-versus-taxonomy trend are support, and the 0.66 is a sanity check, not proof. What remains open is how much the model gives up where the inference stage is wrong, and whether it transfers beyond compact genomes.
Alex: That is a fair summary. If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Sam: Thanks for listening.