Robust Estimation for Dependent Binary Network Data
Tianyu Liu, Somabha Mukherjee, Abhik Ghosh
TL;DR
The paper develops a robust parameter estimation framework for dependent binary network data modeled by Ising/MRFs, using the Minimum Density Power Divergence to generalize the MPL estimator from a single contaminated observation. It establishes $\sqrt{N}$-consistency and central limit theorems for a broad class of $Z$-estimators, including the MDPD, under flexible graph-structure conditions, while proving bounded influence and reduced gross error sensitivity for robustness. Through extensive simulations on Ising models across lattices, ER/ SBM/ SK graphs, and slightly dense graphs, the approach shows improved performance under contamination with little to no loss in clean data efficiency. Real-data analyses in social networks, neurobiology, and genomics demonstrate practical gains in prediction accuracy when using higher $\lambda$, validating the method’s applicability to noisy network data. Overall, the work provides a theoretically grounded, computationally tractable robust alternative to MPL for networked binary data and highlights promising directions for broadening robust network inference.
Abstract
We consider the problem of learning the interaction strength between the nodes of a network based on dependent binary observations residing on these nodes, generated from a Markov Random Field (MRF). Since these observations can possibly be corrupted/noisy in larger networks in practice, it is important to robustly estimate the parameters of the underlying true MRF to account for such inherent contamination in observed data. However, it is well-known that classical likelihood and pseudolikelihood based approaches are highly sensitive to even a small amount of data contamination. So, in this paper, we propose a density power divergence (DPD) based robust generalization of the computationally efficient maximum pseudolikelihood (MPL) estimator of the interaction strength parameter, and derive its rate of consistency under the pure model. Along the way, we establish consistency and asymptotics for a class of general $Z$-estimators, covering our proposed DPD based estimators, under flexible assumptions that hold for a substantial class of standard models. To the best of our knowledge, these are the first central limit theorems for the class of general $Z$-estimators in such settings. Moreover, we show that the gross error sensitivities of the proposed DPD based estimators are significantly smaller than that of the MPL estimator, thereby theoretically justifying the greater (local) robustness of the former under contaminated settings. Finally, we demonstrate the superior (finite sample) performance of the DPD based variants over the traditional MPL estimator in a number of synthetically generated contaminated network datasets, and apply them to learn the network interaction strength in several real datasets from diverse domains of social science, neurobiology and genomics.
