This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility: An international team of researchers, including those at the Earth-Life Science Institute (ELSI) at the Institute of Science Tokyo, has developed a protein language model that brings together two fundamental sources of information about proteins: their amino acid sequences and three-dimensional structures. The model provides researchers with a new way to map relationships across the protein universe and investigate how proteins have evolved over billions of years.
The research was led by Professor Rachel Kolodny and Ph.D. candidate Guy Yanai of the University of Haifa, Professor Nir Ben-Tal and graduate student Gabriel Axel of Tel Aviv University, and Specially Appointed Associate Professor Liam M. Kolodny also spent five months as a visiting researcher at ELSI developing approaches to analyze the new model. The findings are published in Proceedings of the National Academy of Sciences.
Thousands of protein families are responsible for carrying out nearly every function within living cells. A fundamental question in evolutionary biochemistry is how these proteins are related to one another and where they came from in the first place. Scientists traditionally organize proteins into hierarchical groups based on their relatedness, somewhat like the genus and species classifications used for living organisms.
These carefully curated systems contain decades of scientific knowledge, but advances in artificial intelligence are creating new ways to explore relationships across the vast protein universe. Protein language models can convert a protein into a numerical representation known as an "embedding." One way to think of an embedding is as a kind of ZIP code: Proteins with similar properties tend to receive nearby addresses. Researchers can then visualize these relationships to produce a "protein world map." However, there is a complication.
Proteins contain information in both their amino acid sequences and their three-dimensional structures, and the relationship between the two is not straightforward. Proteins with unrelated sequences can sometimes adopt similar structures, while similar or even identical sequences can produce very different structures. Most protein language models have approached protein sequence and structure separately.
Even models that use both kinds of information do not necessarily place the sequence and structure of the same protein at the same location on a protein map. The researchers developed a model called Contrastive Learning Sequence-Structure, or CLSS, designed to produce highly similar embeddings for both the sequence and structure representations of a protein. CLSS uses an approach called contrastive learning.
Extract — continue reading at the source.