Table of Contents
Fetching ...

Reliable data clustering with Bayesian community detection

Magnus Neuman, Jelena Smiljanić, Martin Rosvall

TL;DR

The paper addresses the challenge of clustering noisy, high-dimensional similarity data by eliminating arbitrary sparsification decisions. It introduces a one-step Bayesian framework based on Minimum Description Length (MDL) that unifies sparsification and clustering via the Regularized Map Equation and the Degree-Corrected SBM, using description-length compression $\Delta D(\tau)$ to select the optimal threshold $\tau^*$. In synthetic tests, the Regularized Map Equation reliably recovers planted partitions under high noise and limited samples, outperforming traditional methods and cross-validation; on gene co-expression data it yields more robust, functionally coherent modules than WGCNA and demonstrates greater stability under subsampling. The approach generalizes to various similarity measures and data-scarce domains, offering a principled, data-efficient path to uncover modular structure across neuroscience, genomics, and ecology, without imposing arbitrary thresholds or risking overfitting.

Abstract

From neuroscience and genomics to systems biology and ecology, researchers rely on clustering similarity data to uncover modular structure. Yet widely used clustering methods, such as hierarchical clustering, k-means, and WGCNA, lack principled model selection, leaving them susceptible to noise. A common workaround sparsifies a correlation matrix representation to remove noise before clustering, but this extra step introduces arbitrary thresholds that can distort the structure and lead to unreliable results. To detect reliable clusters, we capitalize on recent advances in network science to unite sparsification and clustering with principled model selection. We test two Bayesian community detection methods, the Degree-Corrected Stochastic Block Model and the Regularized Map Equation, both grounded in the Minimum Description Length principle for model selection. In synthetic data, they outperform traditional approaches, detecting planted clusters under high-noise conditions and with fewer samples. Compared to WGCNA on gene co-expression data, the Regularized Map Equation identifies more robust and functionally coherent gene modules. Our results establish Bayesian community detection as a principled and noise-resistant framework for uncovering modular structure in high-dimensional data across fields.

Reliable data clustering with Bayesian community detection

TL;DR

The paper addresses the challenge of clustering noisy, high-dimensional similarity data by eliminating arbitrary sparsification decisions. It introduces a one-step Bayesian framework based on Minimum Description Length (MDL) that unifies sparsification and clustering via the Regularized Map Equation and the Degree-Corrected SBM, using description-length compression to select the optimal threshold . In synthetic tests, the Regularized Map Equation reliably recovers planted partitions under high noise and limited samples, outperforming traditional methods and cross-validation; on gene co-expression data it yields more robust, functionally coherent modules than WGCNA and demonstrates greater stability under subsampling. The approach generalizes to various similarity measures and data-scarce domains, offering a principled, data-efficient path to uncover modular structure across neuroscience, genomics, and ecology, without imposing arbitrary thresholds or risking overfitting.

Abstract

From neuroscience and genomics to systems biology and ecology, researchers rely on clustering similarity data to uncover modular structure. Yet widely used clustering methods, such as hierarchical clustering, k-means, and WGCNA, lack principled model selection, leaving them susceptible to noise. A common workaround sparsifies a correlation matrix representation to remove noise before clustering, but this extra step introduces arbitrary thresholds that can distort the structure and lead to unreliable results. To detect reliable clusters, we capitalize on recent advances in network science to unite sparsification and clustering with principled model selection. We test two Bayesian community detection methods, the Degree-Corrected Stochastic Block Model and the Regularized Map Equation, both grounded in the Minimum Description Length principle for model selection. In synthetic data, they outperform traditional approaches, detecting planted clusters under high-noise conditions and with fewer samples. Compared to WGCNA on gene co-expression data, the Regularized Map Equation identifies more robust and functionally coherent gene modules. Our results establish Bayesian community detection as a principled and noise-resistant framework for uncovering modular structure in high-dimensional data across fields.
Paper Structure (10 sections, 17 equations, 10 figures)

This paper contains 10 sections, 17 equations, 10 figures.

Figures (10)

  • Figure 1: Overlapping correlation distributions. In A, the planted structure is visible in the observed correlation matrix $\hat{\Sigma}$, since the within-cluster correlations on average exceed the outside-cluster correlations. This separation depends on the number of samples, nodes, and clusters, and on the population correlation $\rho$. Colorbar truncated at 0.4; diagonal entries have correlation 1. With few samples in B, the within- and outside-cluster distributions $\hat{f}$ and $\hat{f}_0$ overlap, and the number of false positives using the threshold $r_i$, where $\hat{f}_0$ and $\hat{f}$ intersect, is relatively large, making correct inference difficult. With more samples in C, the situation becomes more favorable. The fractions of within- and outside-cluster correlations above $r_i$ are $p_{\mathrm{in}}$ (green, true positives) and $p_{\mathrm{out}}$ (orange, false positives) respectively.
  • Figure 2: The transition to separated correlation distributions. In A, the ratio $p_{\mathrm{out}}/p_{\mathrm{in}}$ shows the non-linear transition from overlapping to separated outside- and within-cluster distributions as the number of samples $L$ increases. More samples are required if the number of clusters $q$ increases, particularly if the population correlation is weak. The number of nodes has a negligible effect, as shown by the barely visible dashed lines where the number of nodes is doubled. In B, splitting the network into more clusters increases $p_{\mathrm{out}}/p_{\mathrm{in}}$, making correct inference more difficult, particularly with weak correlation. The number of features matters weakly in this case, where more features actually ease inference.
  • Figure 3: Bayesian community detection and description length compression. In A, the description length compression peaks to the right of the intersection $r_i$ of within- and outside-cluster correlation distributions. The compression peak lies between the dense (B) and sparse (D) networks, where the signal of modular structure is maximized (C). To avoid overfitting modular structure in sparse networks, Bayesian community detection uses a prior shown as orange, thin links in B-D. This approach regularizes correlation networks without splitting already scarce data. The node border colors in network D show the clusters without the prior.
  • Figure 4: Method performance on synthetic data. Performance measured by the adjusted mutual information (AMI) between planted and inferred partitions. For networks with $N=1000$ nodes, correlation strength $\rho=0.2$, and $q=2$ clusters, the noise level is low and all methods but the DC-SBM correctly infer the planted partition using few samples (A). The DC-SBM suffers from assuming link independence while correlation networks tend to close triangles. With $q=50$ clusters, the Regularized Map Equation and the DC-SBM recover the planted partition with fewer samples than cross-validation and hierarchical clustering (B). With $q=50$ clusters, the noise level is higher, and hierarchical clustering struggles to infer the correct partition, as does the DC-SBM due to its resolution limit (C). Hierarchical clustering and the DC-SBM infer clusters (AMI$>$0) even in pure noise (B and C with few samples), making interpretation difficult. For schematic networks with $L=100$ samples, correlation strength $\rho=0.3$, and two 50-node clusters (E, F) or ten 5-node clusters (E, G), the DC-SBM overfits the large clusters due to high clustering coefficients (F) and underfits the small clusters due to its resolution limit (G).
  • Figure 5: Detectability limits in the $L-\rho$ space. Four methods compared on network with $N=1000$ nodes and varying number of clusters $q$. Detectability defined as the line where the AMI between planted and detected partitions on average exceeds 0.95. The link ratio $p_{\mathrm{out}}/p_{\mathrm{in}}$ as contour levels in the background. Cross-validation and hierarchical clustering require relatively strong signals for correct inference, and hierarchical clustering is particularly sensitive to small sample sizes. The DC-SBM works well in the intermediate case with $q=50$ clusters (B), where modules are neither too small nor too large and dense, but otherwise struggles with over- or underfitting. The Regularized Map Equation (RegMapEq) pushes the detectability limit furthest into the high-noise regions across all noise levels and modular structures.
  • ...and 5 more figures