sözaltı news Science
Science
EN AZ
New clustering method uncovers hidden regularities in data

New clustering method uncovers hidden regularities in data

phys.org 02.09.2026 17:00 2 views
One dataset, one model? This approach does not always produce the most useful insights, as a dataset often contains many different relationships. To better understand them, researchers at the Paluno Research Institute in

This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility: One dataset, one model? This approach does not always produce the most useful insights, as a dataset often contains many different relationships.

To better understand them, researchers at the Paluno Research Institute in the Faculty of Computer Science at the University of Duisburg-Essen have developed a method that clusters data according to mathematical functions. The key feature is that neither the number and structure of the functions nor the assignment of data points to them needs to be known in advance. Anyone seeking to draw conclusions about technical, biological or economic systems from measurement data generally has to account for different system behaviors.

These cannot always be described accurately and comprehensibly by a single model. For example, if one part of a dataset follows the function f(x) = 2*x, while another is better described by g(x) = x*x, examining the two parts separately may provide more insight. But how can data be grouped meaningfully when the underlying relationships are not yet known?

One possible solution is to use conventional clustering methods such as K-means, which group data points according to their proximity in feature space. However, if the different behaviors cannot be clearly separated in this space, these methods can fail to identify meaningful groups. This is where the work of Peter Zdankin, Arne Kummerow and Professor Dr.

Torben Weis comes in: Their method, CluBS (Clustering Behavioral Similarity), groups data points according to whether they can be described by the same mathematical equation. The findings are published in the Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2. This method takes a step-by-step approach to finding the result.

First, symbolic regression is used to search for a function that best describes a preliminary set of data points. The form of this function is not specified in advance. Instead, it combines variables, numbers and arithmetic operations to form an equation that fits the data as closely as possible.

Extract — continue reading at the source.

Read full story