Table of Contents
Fetching ...

Composite Lp-quantile regression, near quantile regression and the oracle model selection theory

Fuming Lin

TL;DR

This work introduces composite $L^p$-quantile regression (CLpQR) for high-dimensional settings with heavy-tailed errors by requiring only a finite $2(p-1)$-th moment, and examines its oracle-model selection properties. It establishes the asymptotic normality of CLpQR, derives an explicit asymptotic relative efficiency against least squares and composite quantile regression, and demonstrates that CLpQR can be substantially more efficient under heavy tails. A second contribution, near quantile regression, provides asymptotic normality as $p\to1^+$ and $T\to\infty$, along with a practical smoothing interpretation and a parametric covariance estimator. The paper also develops a unified, efficient algorithm combining cyclic coordinate descent and augmented proximal gradient to fit high-dimensional $L^p$-quantile models, and supports the theory with simulations and a real data application. Overall, the approach offers robust, efficient, and scalable quantile-type modeling for heavy-tailed data with practical guidance on algorithmic implementation and variance estimation.

Abstract

In this paper, we consider high-dimensional Lp-quantile regression which only requires a low order moment of the error and is also a natural generalization of the above methods and Lp-regression as well. The loss function of Lp-quantile regression circumvents the non-differentiability of the absolute loss function and the difficulty of the squares loss function requiring the finiteness of error's variance and thus promises excellent properties of Lp-quantile regression. Specifically, we first develop a new method called composite Lp-quantile regression(CLpQR). We study the oracle model selection theory based on CLpQR (call the estimator CLpQR-oracle) and show in some cases of p CLpQR-oracle behaves better than CQR-oracle (based on composite quantile regression) when error's variance is infinite. Moreover, CLpQR has high efficiency and can be sometimes arbitrarily more efficient than both CQR and the least squares regression. Second, we propose another new regression method,i.e. near quantile regression and prove the asymptotic normality of the estimator when p converges to 1 and the sample size infinity simultaneously. As its applications, a new thought of smoothing quantile objective functions and a new estimation are provided for the asymptotic covariance matrix of quantile regression. Third, we develop a unified efficient algorithm for fitting high-dimensional Lp-quantile regression by combining the cyclic coordinate descent and an augmented proximal gradient algorithm. Remarkably, the algorithm turns out to be a favourable alternative of the commonly used liner programming and interior point algorithm when fitting quantile regression.

Composite Lp-quantile regression, near quantile regression and the oracle model selection theory

TL;DR

This work introduces composite -quantile regression (CLpQR) for high-dimensional settings with heavy-tailed errors by requiring only a finite -th moment, and examines its oracle-model selection properties. It establishes the asymptotic normality of CLpQR, derives an explicit asymptotic relative efficiency against least squares and composite quantile regression, and demonstrates that CLpQR can be substantially more efficient under heavy tails. A second contribution, near quantile regression, provides asymptotic normality as and , along with a practical smoothing interpretation and a parametric covariance estimator. The paper also develops a unified, efficient algorithm combining cyclic coordinate descent and augmented proximal gradient to fit high-dimensional -quantile models, and supports the theory with simulations and a real data application. Overall, the approach offers robust, efficient, and scalable quantile-type modeling for heavy-tailed data with practical guidance on algorithmic implementation and variance estimation.

Abstract

In this paper, we consider high-dimensional Lp-quantile regression which only requires a low order moment of the error and is also a natural generalization of the above methods and Lp-regression as well. The loss function of Lp-quantile regression circumvents the non-differentiability of the absolute loss function and the difficulty of the squares loss function requiring the finiteness of error's variance and thus promises excellent properties of Lp-quantile regression. Specifically, we first develop a new method called composite Lp-quantile regression(CLpQR). We study the oracle model selection theory based on CLpQR (call the estimator CLpQR-oracle) and show in some cases of p CLpQR-oracle behaves better than CQR-oracle (based on composite quantile regression) when error's variance is infinite. Moreover, CLpQR has high efficiency and can be sometimes arbitrarily more efficient than both CQR and the least squares regression. Second, we propose another new regression method,i.e. near quantile regression and prove the asymptotic normality of the estimator when p converges to 1 and the sample size infinity simultaneously. As its applications, a new thought of smoothing quantile objective functions and a new estimation are provided for the asymptotic covariance matrix of quantile regression. Third, we develop a unified efficient algorithm for fitting high-dimensional Lp-quantile regression by combining the cyclic coordinate descent and an augmented proximal gradient algorithm. Remarkably, the algorithm turns out to be a favourable alternative of the commonly used liner programming and interior point algorithm when fitting quantile regression.
Paper Structure (11 sections, 10 theorems, 100 equations, 1 figure, 2 tables)

This paper contains 11 sections, 10 theorems, 100 equations, 1 figure, 2 tables.

Key Result

Theorem 2.1

Suppose $1<p\leq2$ and Assumptions ass2.1-ass2.3 hold, then $\sqrt{T}(\hat{\boldsymbol{\beta}}^{clp}-\boldsymbol{\beta}^{*})$ is asymptotically normal with mean 0 and covariance matrix where $\boldsymbol{\varphi}_{\tau,p}(s)=p|\tau-I(s<0)||s|^{p-1}\hbox{sign}(s)$ and $\boldsymbol{\psi}_{\tau,p}(s)=p(p-1)|\tau-I(s<0)||s|^{p-2}$.

Figures (1)

  • Figure 1: Upper left panel: $ARE_{CQR}$ and $ARE_{CLpQR}$ ($p=1.2$ and 1.2) as the functions of $\rho$ the mixture parameter of the mixture of two normals. Upper right panel: $ARE_{CLpQR}$ ($p\geq1.3$) as the functions of $\rho$. Lower left panel: $ARE_{CLpQR}$ ($\rho=0.9$ and 1) as the functions of $p$. When $p=1$$ARE_{CLpQR}$ is just $ARE_{CQR}$. Lower right panel: $ARE_{CLpQR}$ as the function of $p$ when the error obeys the GED. The horizontal line marked by triangular indicates $ARE_{CQR}=0.8748277$ always.

Theorems & Definitions (15)

  • Remark 2.1
  • Theorem 2.1
  • Theorem 2.2
  • Theorem 3.1
  • Remark 3.1
  • Theorem 4.1
  • Remark 4.1
  • Corollary 4.1
  • Theorem 4.2
  • Remark 5.1
  • ...and 5 more