Table of Contents
Fetching ...

Robust Estimation for Dependent Binary Network Data

Tianyu Liu, Somabha Mukherjee, Abhik Ghosh

TL;DR

The paper develops a robust parameter estimation framework for dependent binary network data modeled by Ising/MRFs, using the Minimum Density Power Divergence to generalize the MPL estimator from a single contaminated observation. It establishes $\sqrt{N}$-consistency and central limit theorems for a broad class of $Z$-estimators, including the MDPD, under flexible graph-structure conditions, while proving bounded influence and reduced gross error sensitivity for robustness. Through extensive simulations on Ising models across lattices, ER/ SBM/ SK graphs, and slightly dense graphs, the approach shows improved performance under contamination with little to no loss in clean data efficiency. Real-data analyses in social networks, neurobiology, and genomics demonstrate practical gains in prediction accuracy when using higher $\lambda$, validating the method’s applicability to noisy network data. Overall, the work provides a theoretically grounded, computationally tractable robust alternative to MPL for networked binary data and highlights promising directions for broadening robust network inference.

Abstract

We consider the problem of learning the interaction strength between the nodes of a network based on dependent binary observations residing on these nodes, generated from a Markov Random Field (MRF). Since these observations can possibly be corrupted/noisy in larger networks in practice, it is important to robustly estimate the parameters of the underlying true MRF to account for such inherent contamination in observed data. However, it is well-known that classical likelihood and pseudolikelihood based approaches are highly sensitive to even a small amount of data contamination. So, in this paper, we propose a density power divergence (DPD) based robust generalization of the computationally efficient maximum pseudolikelihood (MPL) estimator of the interaction strength parameter, and derive its rate of consistency under the pure model. Along the way, we establish consistency and asymptotics for a class of general $Z$-estimators, covering our proposed DPD based estimators, under flexible assumptions that hold for a substantial class of standard models. To the best of our knowledge, these are the first central limit theorems for the class of general $Z$-estimators in such settings. Moreover, we show that the gross error sensitivities of the proposed DPD based estimators are significantly smaller than that of the MPL estimator, thereby theoretically justifying the greater (local) robustness of the former under contaminated settings. Finally, we demonstrate the superior (finite sample) performance of the DPD based variants over the traditional MPL estimator in a number of synthetically generated contaminated network datasets, and apply them to learn the network interaction strength in several real datasets from diverse domains of social science, neurobiology and genomics.

Robust Estimation for Dependent Binary Network Data

TL;DR

The paper develops a robust parameter estimation framework for dependent binary network data modeled by Ising/MRFs, using the Minimum Density Power Divergence to generalize the MPL estimator from a single contaminated observation. It establishes -consistency and central limit theorems for a broad class of -estimators, including the MDPD, under flexible graph-structure conditions, while proving bounded influence and reduced gross error sensitivity for robustness. Through extensive simulations on Ising models across lattices, ER/ SBM/ SK graphs, and slightly dense graphs, the approach shows improved performance under contamination with little to no loss in clean data efficiency. Real-data analyses in social networks, neurobiology, and genomics demonstrate practical gains in prediction accuracy when using higher , validating the method’s applicability to noisy network data. Overall, the work provides a theoretically grounded, computationally tractable robust alternative to MPL for networked binary data and highlights promising directions for broadening robust network inference.

Abstract

We consider the problem of learning the interaction strength between the nodes of a network based on dependent binary observations residing on these nodes, generated from a Markov Random Field (MRF). Since these observations can possibly be corrupted/noisy in larger networks in practice, it is important to robustly estimate the parameters of the underlying true MRF to account for such inherent contamination in observed data. However, it is well-known that classical likelihood and pseudolikelihood based approaches are highly sensitive to even a small amount of data contamination. So, in this paper, we propose a density power divergence (DPD) based robust generalization of the computationally efficient maximum pseudolikelihood (MPL) estimator of the interaction strength parameter, and derive its rate of consistency under the pure model. Along the way, we establish consistency and asymptotics for a class of general -estimators, covering our proposed DPD based estimators, under flexible assumptions that hold for a substantial class of standard models. To the best of our knowledge, these are the first central limit theorems for the class of general -estimators in such settings. Moreover, we show that the gross error sensitivities of the proposed DPD based estimators are significantly smaller than that of the MPL estimator, thereby theoretically justifying the greater (local) robustness of the former under contaminated settings. Finally, we demonstrate the superior (finite sample) performance of the DPD based variants over the traditional MPL estimator in a number of synthetically generated contaminated network datasets, and apply them to learn the network interaction strength in several real datasets from diverse domains of social science, neurobiology and genomics.
Paper Structure (25 sections, 10 theorems, 149 equations, 14 figures, 1 table)

This paper contains 25 sections, 10 theorems, 149 equations, 14 figures, 1 table.

Key Result

Theorem 1

Let $\hat{\beta}_{\lambda}$ be the MDPD estimator of $\beta$ for some fixed $\lambda>0$, based on a single observation $\bm X$ from the Ising model modeldef. Suppose that the following conditions hold: where $\|\cdot\|$ denotes the spectral norm. Then $\hat{\beta}_{\lambda}$ is a $\sqrt{N}$-consistent sequence of estimators for $\beta$, i.e. $\sqrt{N}(\hat{\beta}_{\lambda}-\beta) = O_P(1)$.

Figures (14)

  • Figure 1: The GES for MDPD estimates from Ising models on different network structures, obtained numerically at $N=100$.
  • Figure 2: MSE and bias for estimates from 1-D lattice, with $\beta = 0.5$, $N=2000$ and contamination levels set to $0\%$, $20\%$ and $40\%$.
  • Figure 3: MSE and bias for estimates from 2-D lattice, with $\beta = 0.5$, $N=6400$ and contamination levels set to $0\%$, $20\%$ and $40\%$.
  • Figure 4: MSE and bias for estimates from ErdÅ‘s–Rényi random graph with $p=\frac{5}{N}$, $\beta = 0.8$ and $N=2000$.
  • Figure 5: MSE and bias for estimates from ErdÅ‘s–Rényi random graph with varying $\beta$, $p=\frac{5}{N}$ and $N=2000$.
  • ...and 9 more figures

Theorems & Definitions (18)

  • Theorem 1
  • Example 1: Non mean-field interactions/bounded-degree graphs
  • Example 2: Regular graphs
  • Example 3: Sequence of dense graphs converging to a Graphon
  • Example 4: Spin glass models
  • Theorem 2
  • Theorem 3
  • Definition 1
  • Theorem 4
  • Lemma 1
  • ...and 8 more