This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility: A research team at the University of Vienna, led by pharmaceutical chemist Johannes Kirchmair, has developed a statistical method to identify hidden anomalies in scientific datasets. The method can be applied automatically to large datasets, making it easier for researchers to pinpoint those that need closer scrutiny.
The approach improves the quality of data-driven research and supports the development of reliable AI applications. Modern scientific methods increasingly rely on large volumes of data. In drug discovery, for example, millions of experimental measurements are often collected from databases and fed into AI models to aid the discovery and development of new medicines.
Data quality is crucial: Errors in data collection, processing or integration can slow down or even undermine subsequent analyses, findings and research. But systematically checking large datasets poses considerable challenges for researchers. Uday Abu-Shehab, Matthias Welsch and Kirchmair, from the Christian Doppler Laboratory for Molecular Informatics in the Biosciences at the University of Vienna's Department of Pharmaceutical Sciences, developed a method to automatically identify anomalous datasets.
At its heart is Benford's law, which describes the observation that, in many naturally occurring datasets, certain digits appear as the first digit more often than others. A substantial deviation from this pattern may indicate irregularities or potential quality issues. The findings are published in the journal Patterns.
"Our method does not provide direct evidence that data is flawed or falsified," Kirchmair explains. "Rather, it helps to filter out datasets from large collections that should be subjected to closer scrutiny. This enables researchers to focus their attention specifically on potentially problematic datasets." To help researchers prioritize datasets, the team developed the "Simulation-Based Benfordness Estimation" (SBBE) method.
It combines extensive computer simulations with Bayesian statistics to provide both an assessment of data quality and an estimate of the associated uncertainty. This makes it possible to compare datasets more effectively than with previous methods. "With many existing approaches, the results of quality analyses depend heavily on the number of available data points," Abu-Shehab explains.
Extract — continue reading at the source.