Table of Contents
Fetching ...

Measuring nodes centrality when local and global measures overlap

Lorenzo Costantini, Carla Sciarra, Luca Ridolfi, Francesco Laio

TL;DR

The paper addresses the limitation that in networks with a high spectral gap, global centrality metrics largely replicate degree, hindering discovery of new node roles. It introduces GENEPY, a two-eigenvector generalized economic complexity index applied to the degree-filtered proximity matrices $N$ and $G$, to capture connectivity-pattern centrality beyond degree. Through synthetic PTN and BERG networks and on 284 real networks, GENEPY provides centrality rankings that are less collinear with degree or eigenvector and reveals complementary information about node importance, including anti-centrality behavior on one side of bipartite mappings. The approach is applicable to both bipartite and mapped monopartite networks and is supported by a Variance Inflation Factor analysis showing non-collinearity. The results suggest GENEPY as a practical tool for unveiling nuanced centrality in high spectral gap systems and point to future predictive comparisons and broader network class extensions.

Abstract

Centrality metrics aim to identify the most relevant nodes in a network. In literature, a broad set of metrics exists, either measuring local or global centrality characteristics. Nevertheless, when networks exhibit a high spectral gap, the usual global centrality measures typically do not add significant information with respect to the degree, i.e., the simplest local metric. To extract new information from this class of networks, we propose the use of the GENeralized Economic comPlexitY index (GENEPY). Despite its original definition within the economic field, the GENEPY can be easily applied and interpreted on a wide range of networks, characterized by high spectral gap, including monopartite and bipartite networks systems. Tests on synthetic and real-world networks show that the GENEPY can shed new light about the nodes centrality, carrying information generally poorly correlated with the nodes number of direct connections (nodes degree).

Measuring nodes centrality when local and global measures overlap

TL;DR

The paper addresses the limitation that in networks with a high spectral gap, global centrality metrics largely replicate degree, hindering discovery of new node roles. It introduces GENEPY, a two-eigenvector generalized economic complexity index applied to the degree-filtered proximity matrices and , to capture connectivity-pattern centrality beyond degree. Through synthetic PTN and BERG networks and on 284 real networks, GENEPY provides centrality rankings that are less collinear with degree or eigenvector and reveals complementary information about node importance, including anti-centrality behavior on one side of bipartite mappings. The approach is applicable to both bipartite and mapped monopartite networks and is supported by a Variance Inflation Factor analysis showing non-collinearity. The results suggest GENEPY as a practical tool for unveiling nuanced centrality in high spectral gap systems and point to future predictive comparisons and broader network class extensions.

Abstract

Centrality metrics aim to identify the most relevant nodes in a network. In literature, a broad set of metrics exists, either measuring local or global centrality characteristics. Nevertheless, when networks exhibit a high spectral gap, the usual global centrality measures typically do not add significant information with respect to the degree, i.e., the simplest local metric. To extract new information from this class of networks, we propose the use of the GENeralized Economic comPlexitY index (GENEPY). Despite its original definition within the economic field, the GENEPY can be easily applied and interpreted on a wide range of networks, characterized by high spectral gap, including monopartite and bipartite networks systems. Tests on synthetic and real-world networks show that the GENEPY can shed new light about the nodes centrality, carrying information generally poorly correlated with the nodes number of direct connections (nodes degree).
Paper Structure (10 sections, 8 equations, 5 figures, 2 tables)

This paper contains 10 sections, 8 equations, 5 figures, 2 tables.

Figures (5)

  • Figure 1: Examples of two incidence matrices for each class of artificially generated networks of dimension $40 \times 250$ with a specific value of the parameter $\beta$ (yellow entries are nonzero values and blue otherwise) and corresponding correlation values between the degree and other centrality metrics as applied to the set of equally sized networks. The Spearman’s correlation coefficients ($\rho_S$) between the degree (D) and the eigenvector (E), closeness (C), betweenness (B), subgraph (SG) centrality, and total communicability (TC) are defined as functions of the relative spectral gap, and the computation of the metrics for both nodes in the $U$ and $P$ sets of the networks is detailed. Top panels refer to Pseudo-Triangular Networks (PTNs), while bottom ones to Bipartite Erdős-Rényi Graphs (BERGs). In panels (b), (c), (e) and (f), each point represents the mean among $N_{sim} = 100$ networks of the considered class, the whiskers and the shaded regions describe $\pm1$ standard deviation of the correlation values and relative spectral gap, respectively. The arrows indicate the direction in which the threshold values $\beta$ increase for both PTNs and BERGs.
  • Figure 2: The Spearman's correlation coefficients among different centrality measures for the Pseudo-Triangular Networks (PTNs), panels (a) and (b), and BERGs ones (Bipartite Erdős-Rényi Graphs), panels (c) and (d). The correlation values ($\rho_S$) between the degree (D) and eigenvector (E) centrality (blue dots), and the degree and GENEPY index (red squares), are shown as functions of the relative spectral gap. The right panels refer to the $U$ set, the left ones to the $P$ set. Each point represents the average value of $N_{sim}=100$ realizations, and $\pm1$ standard deviations are given for both correlation (whiskers) and relative spectral gap (shaded regions) values. In purple, we show the mean Variance Inflation Factor (VIF) computed over the $N_{sim}$ artificial networks between the degree and eigenvector centrality (diamonds) and the degree-GENEPY values (stars).
  • Figure 3: Values of the Variance Inflation Factor computed among the centrality metrics (degree -- D --, eigenvector -- E, blue circles --, closeness -- C, red squares --, betweenness -- B, yellow triangles --, subgraph centrality -- SG, purple pentagrams -- and total communicability -- TC, green diamonds --) applied onto the Women-Events network TwoModeDatadavis2009deep and its simulation with edges' addition/removal. The benchmark for the computation of the VIF values is the degree centrality in the top-panels (panel (a) for the women set, and panel (b) for the events set); the GENEPY index computed onto N in panel (c); the GENEPY index computed onto G in panel (d). To modify the spectral gap, $100$ ($50$) links were randomly added (removed), one at a time, to the original network. Each point is the mean, computed at the same number of links in the network, among $100$ repetition of the edges' addition/removal process. For the sake of graphical representation, the infinite values of the VIF (corresponding to a Spearman's correlation of $1$) were replaced by VIF$=1000$ (Spearman's correlation value $0.9995$). The black arrows indicate the direction in which the links were added, and the vertical line in each panel highlights the VIF values and relative spectral gap of the original network.
  • Figure 4: Comparison of the VIF values between the GENEPY and the degree (D, panel (a)), eigenvector (E, panel (b)), closeness (C, panel (c)), betweenness (B, panel (d)), subgraph (SG, panel (e)) and total communicability (TC, panel (f)) centrality measures as a function of the relative spectral gap for the $284$ real-world networks considered in this work. In each panel, the blue circles (red diamond) report the VIF values computed between the GENEPY calculated onto the proximity matrix N (G) with the Centrality Measure (CM) reported in the panel title. The thick blue dashed (thick red solid) lines fit the logarithm (base $10$) of blue circles (red diamonds) with a function $a(SG_r)^b+c$. The thin dashed and solid black lines describe the fits of VIF values between the degree centrality and the centrality measure reported in the title of each panel. For sake of graphical representation, the y-axis is limited to the range $[0, 100]$ to enhance the readability of the points.
  • Figure 5: Representation of the bipartite network used as toy model. The nodes in the set $U$ (indicated with the Greek letters $\alpha$, $\beta$, $\gamma$, $\delta$ and $\epsilon$) are represented by red squares, while those in $P$ (lowercase Latin letters $a$, $b$, $c$, $d$, $e$ and $f$) by blue circles. The black solid lines represent the connections among the nodes in the two sets. The nodes in $U$ are ranked according to the GENEPY computed onto N, while those in $P$ onto the proximity matrix G.