Table of Contents
Fetching ...

Beyond False Discovery Rate: A Stepdown Group SLOPE Approach for Grouped Variable Selection

Xuelin Zhang, Jingxuan Liang, Xinyue Liu, Hong Chen, Biqin Song

TL;DR

The Group Stepdown SLOPE is introduced, a unified optimization procedure which is capable of embedding the Lehmann-Romano stepdown rules into SLOPE to achieve finite-sample guarantees under k-FWER and FDP thresholds and derive closed-form regularization sequences under orthogonal designs that provably bound k-FWER and FDP at user-specified levels.

Abstract

High-dimensional feature selection is routinely required to balance statistical power with strict control of multiple-error metrics such as the k-Family-Wise Error Rate (k-FWER) and the False Discovery Proportion (FDP), yet some existing frameworks are confined to the narrower goal of controlling the expected False Discovery Rate (FDR) and can not exploit the group-structure of the covariates, such as Sorted L-One Penalized Estimation (SLOPE). We introduce the Group Stepdown SLOPE, a unified optimization procedure which is capable of embedding the Lehmann-Romano stepdown rules into SLOPE to achieve finite-sample guarantees under k-FWER and FDP thresholds. Specifically, we derive closed-form regularization sequences under orthogonal designs that provably bound k-FWER and FDP at user-specified levels, and extend these results to grouped settings via gk-SLOPE and gF-SLOPE, which control the analogous group-level errors gk-FWER and gFDP. For non-orthogonal general designs, we provide a calibrated data-driven sequence inspired by Gaussian approximation and Monte-Carlo correction, preserving convexity and scalability. Extensive simulations are conducted across sparse, correlated, and group-structured regimes. Empirical results corroborate our theoretical findings that the proposed methods achieve nominal error control, while yielding markedly higher power than competing stepdown procedures, thereby confirming the practical value of the theoretical advances.

Beyond False Discovery Rate: A Stepdown Group SLOPE Approach for Grouped Variable Selection

TL;DR

The Group Stepdown SLOPE is introduced, a unified optimization procedure which is capable of embedding the Lehmann-Romano stepdown rules into SLOPE to achieve finite-sample guarantees under k-FWER and FDP thresholds and derive closed-form regularization sequences under orthogonal designs that provably bound k-FWER and FDP at user-specified levels.

Abstract

High-dimensional feature selection is routinely required to balance statistical power with strict control of multiple-error metrics such as the k-Family-Wise Error Rate (k-FWER) and the False Discovery Proportion (FDP), yet some existing frameworks are confined to the narrower goal of controlling the expected False Discovery Rate (FDR) and can not exploit the group-structure of the covariates, such as Sorted L-One Penalized Estimation (SLOPE). We introduce the Group Stepdown SLOPE, a unified optimization procedure which is capable of embedding the Lehmann-Romano stepdown rules into SLOPE to achieve finite-sample guarantees under k-FWER and FDP thresholds. Specifically, we derive closed-form regularization sequences under orthogonal designs that provably bound k-FWER and FDP at user-specified levels, and extend these results to grouped settings via gk-SLOPE and gF-SLOPE, which control the analogous group-level errors gk-FWER and gFDP. For non-orthogonal general designs, we provide a calibrated data-driven sequence inspired by Gaussian approximation and Monte-Carlo correction, preserving convexity and scalability. Extensive simulations are conducted across sparse, correlated, and group-structured regimes. Empirical results corroborate our theoretical findings that the proposed methods achieve nominal error control, while yielding markedly higher power than competing stepdown procedures, thereby confirming the practical value of the theoretical advances.
Paper Structure (19 sections, 10 theorems, 73 equations, 7 figures, 11 tables, 3 algorithms)

This paper contains 19 sections, 10 theorems, 73 equations, 7 figures, 11 tables, 3 algorithms.

Key Result

Lemma 1

slope In the linear model with the orthogonal design $X$ and $\epsilon\sim N (0,\sigma^2 I_n)$, the SLOPE SLOPEeq with the regularization parameter sequence (lambdaBH) satisfies $\mathrm{FDR} \leq \frac{m_{0} q}{m}$, where $m_0$ is the number of true null hypotheses and $q$ is the desired FDR level.

Figures (7)

  • Figure 1: $k$-FWER provided by different approaches for controlled feature selection under orthogonal design (with different $k$ and $t$). The value in the small square is the size of $k$-FWER. The darker the color, the larger the $k$-FWER and vice versa.
  • Figure 2: Result for controlled feature selection on the simulated data. The black dashed lines indicate the target FDR level. Constance for $k$-SLOPE is $k=6$ in the second column (from left to right). The value in the small square is the size of $k$-FWER in the third and fourth columns (from left to right). The darker the color, the larger the $k$-FWER and vice versa.
  • Figure 3: Power and FDR of F-SLOPE under Gaussian design (different $t$). The black dashed line indicates the target FDR level.
  • Figure 4: The g$k$-FWER provided by different approaches for controlled feature selection under orthogonal design. The value in the small square is the size of g$k$-FWER. The darker the color, the larger the g$k$-FWER and vice versa.
  • Figure 5: The gFDR of g-SLOPE, g$k$-SLOPE and gF-SLOPE on the simulated data in orthogonal design when $k=15$.
  • ...and 2 more figures

Theorems & Definitions (19)

  • Lemma 1
  • Lemma 2
  • Lemma 3
  • Lemma 4
  • Remark 1
  • Lemma 5
  • Remark 2
  • Lemma 6
  • Remark 3
  • Theorem 1
  • ...and 9 more