This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility: Patents are far more than legal documents protecting intellectual property. They are a vast source of scientific information about chemicals under development, often years before they appear in products or in the environment.
That makes patents a valuable data source for early warning systems (EWS), which authorities use to identify potentially hazardous chemicals before they become a threat to human health and the environment. There is just one problem: Much of the chemical information in patents is locked inside drawings. Molecular structures and reaction schemes are embedded as images that no text search can read.
Less than 7% of the patents examined in the study contained chemical structure images at all, but when they did, the structures were essentially invisible to computers unless converted into machine-readable formats. "A patent might depict a new molecule years before it is produced at scale," says Farina Tariq, a doctoral researcher at the Department of Chemistry at Umeå University. "If we can teach computers to read those drawings, we can give scientists and authorities an earlier opportunity to investigate potentially concerning chemicals.
But first, we need to know how reliable the AI tools actually are when applied to real patent documents, not just clean images from databases." The potential workflow is straightforward: A patent is published → an AI system detects a chemical structure in an image → the structure is converted into a machine-readable format → the system compares it with known chemicals and chemical classes → potentially concerning or entirely novel structures are flagged for expert investigation. Once a structure is machine-readable, it can be checked against known hazards: Is this chemical already known? Is it structurally related to a known hazardous substance?
Does it contain structural features associated with particular hazards? Or is it an entirely new structure for which little hazard information exists? If so, can we compute its potential hazards?
In the study, published in the journal Chemical Research in Toxicology, the research team tested three widely used chemical structure recognition tools on two specially curated data sets from the European Patent Office's Espacenet database: one of general organic chemistry and one of per- and polyfluoroalkyl substances (PFAS). The AI-generated structures were then validated by five chemistry experts. On standard, well-drawn organic structures, all three tools performed well, with accuracies of around 74% to 78%, correctly decoding aromatic rings, common functional groups and standard abbreviations.
Extract — continue reading at the source.