Decorative Image
© Birgit Kremer, AI-generated

New Clustering Method

Hidden Regularities in Data

  • von Birgit Kremer
  • 02.09.2026

One dataset, one model? This approach does not always produce the most useful insights, as a dataset often contains many different relationships. To better understand them, researchers at the paluno research institute in the Faculty of Computer Science at the University of Duisburg-Essen have developed a method that clusters data according to mathematical functions. The key feature is that neither the number and structure of the functions nor the assignment of data points to them needs to be known in advance.

Anyone seeking to draw conclusions about technical, biological or economic systems from measurement data generally has to account for different system behaviours. These cannot always be described accurately and comprehensibly by a single model. For example, if one part of a dataset follows the function f(x) = 2*x, while another is better described by g(x) = x*x, examining the two parts separately may provide more insight. But how can data be grouped meaningfully when the underlying relationships are not yet known?

One possible solution is to use conventional clustering methods such as K-means, which group data points according to their proximity in feature space. However, if the different behaviours cannot be clearly separated in this space, these methods can fail to identify meaningful groups. This is where the work of Peter Zdankin, Arne Kummerow and Professor Dr. Torben Weis comes in: their method, CluBS (Clustering Behavioural Similarity), groups data points according to whether they can be described by the same mathematical equation.

This method takes a step-by-step approach to finding the result. First, symbolic regression is used to search for a function that best describes a preliminary set of data points. The form of this function is not specified in advance. Instead, it combines variables, numbers, and arithmetic operations to form an equation that fits the data as closely as possible. CluBS then determines which additional data points belong to this “CluB” — in other words, which points can be described by the same function. This alternating process of function discovery and “CluB” assignment continues until the function and the composition of the “CluB” barely change. The process then starts again with the remaining data until an overview of the different behaviours contained in the dataset has been produced.

The research team evaluated CluBS on 211 real-world and synthetic datasets. Particularly in real-world data, the method frequently found evidence of different behaviours that could be described more accurately by multiple functions than by a single “one-size-fits-all” model. The researchers now aim to further develop the method’s potential for predictive applications. Their goal is to enable CluBS to select the appropriate function reliably even when the target values are unknown.

 

Original Publication: Peter Zdankin, Arne Kummerow, and Torben Weis. 2026. CluBS: Clustering Behavioural Similarity. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD '26). Association for Computing Machinery, New York, NY, USA, 6385–6396.
https://doi.org/10.1145/3770855.3817875

Further Information:
Prof. Dr.-Ing. Torben Weis, Faculty of Computer Science, +49 203/37 9-4210, torben.weis@uni-due.de

Editor: Birgit Kremer, paluno, +49 201/18-3 4655,birgit.kremer@paluno.uni-due.de

Zurück