This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility: New York University researchers have trained an AI model to learn chemical patterns associated with stability in drug-like molecules and accurately predict where their hydrogen atoms should be positioned. Their research, published in the journal Chemical Science, addresses a longstanding challenge in molecular design and drug discovery: how to rapidly and reliably determine the stable form of molecules that share a molecular formula but readily convert into different forms.
Many drug-like molecules can exist in two or more closely related forms called tautomers. A hydrogen atom moves from one site to another, accompanied by a change in the bonding pattern. "Although this may seem like a small change, different tautomers of the same molecule can alter how a molecule interacts with a protein target," explained Yingkai Zhang, professor of chemistry at NYU and the study's senior author.
"Correct tautomer assignment is consequently important for molecular modeling and structure-based drug discovery." Yet determining the correct tautomer remains a challenge—like finding a needle in a moving haystack. This is, in part, because of a scarcity of experimental data characterizing tautomer structures. For instance, in the Protein Data Bank, a global repository of the 3D structures of large biological molecules widely used in biological research, the locations of hydrogen atoms that distinguish one tautomer from another are typically not available.
The Protein Data Bank contains molecular structures that are largely determined using experimental methods such as X-ray crystallography, but the resolution of macromolecular X-ray structures is generally insufficient to reliably locate hydrogen atoms. As a result, hydrogen positions—and therefore the tautomeric state of a molecule—often have to be inferred by scientists. Other methods also fall short.
Using quantum mechanics to assess tautomers is often too computationally demanding and expensive for screening full libraries, while machine learning is limited by the relatively small datasets of experimentally characterized tautomers in solution, which tend to contain only a few hundred molecules. In contrast, many high-resolution small-molecule X-ray crystal structures in the Cambridge Structural Database—the largest repository of experimental crystal structures in the world—show the positions of hydrogen atoms. "We realized that experimentally resolved hydrogen positions in high-resolution small-molecule crystal structures provide a largely untapped source of experimental information about tautomer stability," said Zhang, who is also part of the NYU Simons Center for Computational Physical Chemistry.
Xiaolin Pan, a postdoctoral researcher in Zhang's lab and the study's first author, systematically mined the Cambridge Structural Database to construct a dataset containing more than 1.1 million tautomeric states, orders of magnitude larger than existing experimental datasets. He then trained a graph neural network—a type of artificial intelligence that uses deep learning to find patterns between connected data points—to predict stable tautomers directly from 2D molecular forms, without requiring 3D structures or quantum-mechanical calculations. When the researchers applied their model to 5,075 PDBbind ligands—biomolecular complexes found in the Protein Data Bank—with multiple possible tautomeric states, they identified 126 cases—approximately 2.5%—in which the assigned ligand tautomer was likely incorrect.
Extract — continue reading at the source.