Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency without Model Sweeps
Do Tien Hai, Trung Nguyen Mai, TrungTin Nguyen, Nhat Ho, Binh T. Nguyen, Christopher Drovandi
TL;DR
This work develops a fast-rate, geometry-aware framework for softmax-gated Gaussian Mixture of Experts (SGMoE) by introducing a Voronoi-type loss and a hierarchical, dendrogram-based aggregation path. It shows that over-specification creates slow directions when multiple fitted atoms occupy the same Voronoi cell, but repeated, SGMoE-tailored merges restore near-parametric rates and yield a monotone strengthening of the loss along the aggregation path. A novel fast-rate-aware distance, merged-moment penalties, and a height–likelihood dendrogram selector (DSC) enable sweep-free model selection that is consistent and achieves pointwise optimal parameter rates, even under misspecification. The approach is validated by simulations and a maize proteomics dataset, where DSC robustly recovers the true component count, stabilizes likelihood early, and produces interpretable genotype–phenotype maps. Overall, the work provides a principled pathway to accurate estimation and model order selection in SGMoE without costly multi-size training.
Abstract
We develop a unified statistical framework for softmax-gated Gaussian mixture of experts (SGMoE) that addresses three long-standing obstacles in parameter estimation and model selection: (i) non-identifiability of gating parameters up to common translations, (ii) intrinsic gate-expert interactions that induce coupled differential relations in the likelihood, and (iii) the tight numerator-denominator coupling in the softmax-induced conditional density. Our approach introduces Voronoi-type loss functions aligned with the gate-partition geometry and establishes finite-sample convergence rates for the maximum likelihood estimator (MLE). In over-specified models, we reveal a link between the MLE's convergence rate and the solvability of an associated system of polynomial equations characterizing near-nonidentifiable directions. For model selection, we adapt dendrograms of mixing measures to SGMoE, yielding a consistent, sweep-free selector of the number of experts that attains pointwise-optimal parameter rates under overfitting while avoiding multi-size training. Simulations on synthetic data corroborate the theory, accurately recovering the expert count and achieving the predicted rates for parameter estimation while closely approximating the regression function. Under model misspecification (e.g., $ε$-contamination), the dendrogram selection criterion is robust, recovering the true number of mixture components, while the Akaike information criterion, the Bayesian information criterion, and the integrated completed likelihood tend to overselect as sample size grows. On a maize proteomics dataset of drought-responsive traits, our dendrogram-guided SGMoE selects two experts, exposes a clear mixing-measure hierarchy, stabilizes the likelihood early, and yields interpretable genotype-phenotype maps, outperforming standard criteria without multi-size training.
