Berkeley's GPN-Star uses evolutionary signals to flag disease-linked DNA variants
The genomic language model outperforms competitors at identifying variants that shape inherited traits while using less compute than larger models.

Researchers at UC Berkeley have created a genomic language model called GPN-Star that they say far outpaces competitors at identifying the most important genetic variants contributing to inherited traits, including those that lead to disease. Berkeley News reports the model is also far more computationally efficient than larger models. Genomic language models are trained on vast troves of DNA sequences, much as chatbots are trained on text. GPN-Star learns from evolution, drawing on how sequences have been conserved or changed over time to judge which variants matter. The motivation is a long-standing gap in biology. More than two decades after the human genome was fully sequenced, the meaning of much of its 3 billion base pairs remains unclear. An estimated 1 to 2 percent of human DNA codes for proteins. The rest mixes evolutionary holdovers that no longer code for anything with regulatory elements that control when, where, and how strongly genes are expressed. Those non-coding regions could hold the key to understanding inherited traits, including ones tied to cancer, heart disease, and autism, but scientists first need to understand how variants in this DNA contribute to individual differences. A tool that ranks variants by likely importance, and does so cheaply, could help geneticists prioritize which of the many differences between individuals deserve experimental follow-up. The report does not detail the training data or the comparison benchmarks used.