Reliable data clustering with Bayesian community detection
Magnus Neuman, Jelena Smiljanić, Martin Rosvall
TL;DR
The paper addresses the challenge of clustering noisy, high-dimensional similarity data by eliminating arbitrary sparsification decisions. It introduces a one-step Bayesian framework based on Minimum Description Length (MDL) that unifies sparsification and clustering via the Regularized Map Equation and the Degree-Corrected SBM, using description-length compression $\Delta D(\tau)$ to select the optimal threshold $\tau^*$. In synthetic tests, the Regularized Map Equation reliably recovers planted partitions under high noise and limited samples, outperforming traditional methods and cross-validation; on gene co-expression data it yields more robust, functionally coherent modules than WGCNA and demonstrates greater stability under subsampling. The approach generalizes to various similarity measures and data-scarce domains, offering a principled, data-efficient path to uncover modular structure across neuroscience, genomics, and ecology, without imposing arbitrary thresholds or risking overfitting.
Abstract
From neuroscience and genomics to systems biology and ecology, researchers rely on clustering similarity data to uncover modular structure. Yet widely used clustering methods, such as hierarchical clustering, k-means, and WGCNA, lack principled model selection, leaving them susceptible to noise. A common workaround sparsifies a correlation matrix representation to remove noise before clustering, but this extra step introduces arbitrary thresholds that can distort the structure and lead to unreliable results. To detect reliable clusters, we capitalize on recent advances in network science to unite sparsification and clustering with principled model selection. We test two Bayesian community detection methods, the Degree-Corrected Stochastic Block Model and the Regularized Map Equation, both grounded in the Minimum Description Length principle for model selection. In synthetic data, they outperform traditional approaches, detecting planted clusters under high-noise conditions and with fewer samples. Compared to WGCNA on gene co-expression data, the Regularized Map Equation identifies more robust and functionally coherent gene modules. Our results establish Bayesian community detection as a principled and noise-resistant framework for uncovering modular structure in high-dimensional data across fields.
