Finite sample rates for logistic regression with small noise or few samples

Felix Kuchelmeister; Sara van de Geer

Finite sample rates for logistic regression with small noise or few samples

Felix Kuchelmeister, Sara van de Geer

TL;DR

The logistic regression estimator is studied with Gaussian covariates and labels generated by the Gaussian link function, with a mild optimization constraint on the estimator's length to ensure existence, and finite sample guarantees for its direction and Euclidean norm are provided.

Abstract

The logistic regression estimator is known to inflate the magnitude of its coefficients if the sample size $n$ is small, the dimension $p$ is (moderately) large or the signal-to-noise ratio $1/σ$ is large (probabilities of observing a label are close to 0 or 1). With this in mind, we study the logistic regression estimator with $p\ll n/\log n$, assuming Gaussian covariates and labels generated by the Gaussian link function, with a mild optimization constraint on the estimator's length to ensure existence. We provide finite sample guarantees for its direction, which serves as a classifier, and its Euclidean norm, which is an estimator for the signal-to-noise ratio. We distinguish between two regimes. In the low-noise/small-sample regime ($σ\lesssim (p\log n)/n$), we show that the estimator's direction (and consequentially the classification error) achieve the rate $(p\log n)/n$ - up to the log term as if the problem was noiseless. In this case, the norm of the estimator is at least of order $n/(p\log n)$. If instead $(p\log n)/n\lesssim σ\lesssim 1$, the estimator's direction achieves the rate $\sqrt{σp\log n/n}$, whereas its norm converges to the true norm at the rate $\sqrt{p\log n/(nσ^3)}$. As a corollary, the data are not linearly separable with high probability in this regime. In either regime, logistic regression provides a competitive classifier.

Finite sample rates for logistic regression with small noise or few samples

TL;DR

Abstract

The logistic regression estimator is known to inflate the magnitude of its coefficients if the sample size

is small, the dimension

is (moderately) large or the signal-to-noise ratio

is large (probabilities of observing a label are close to 0 or 1). With this in mind, we study the logistic regression estimator with

, assuming Gaussian covariates and labels generated by the Gaussian link function, with a mild optimization constraint on the estimator's length to ensure existence. We provide finite sample guarantees for its direction, which serves as a classifier, and its Euclidean norm, which is an estimator for the signal-to-noise ratio. We distinguish between two regimes. In the low-noise/small-sample regime (

), we show that the estimator's direction (and consequentially the classification error) achieve the rate

- up to the log term as if the problem was noiseless. In this case, the norm of the estimator is at least of order

. If instead

, the estimator's direction achieves the rate

, whereas its norm converges to the true norm at the rate

. As a corollary, the data are not linearly separable with high probability in this regime. In either regime, logistic regression provides a competitive classifier.

Paper Structure (63 sections, 45 theorems, 354 equations, 1 figure)

This paper contains 63 sections, 45 theorems, 354 equations, 1 figure.

Abstract
Introduction
Problems of logistic regression
Linear separation
Overestimating coefficient magnitude
Inadequacy of classical asymptotic approximations
Related results
Outline
Notation
Main results
Regime 1 ("large noise")
Statement of the result
Idea of the proof
Regime 2 ("small noise")
Statement of the result
...and 48 more sections

Key Result

Theorem 2.1.1

For any $t>0$, if then, with probability at least $1-4\exp(-t)$,

Figures (1)

Figure 1: This figure shows the structure of the paper. The main results are given in Section \ref{['sec_main']} and are proved in \ref{['sec_state']}. These proofs rely on Sections \ref{['sec_geo']}, \ref{['sec_con_b']} and \ref{['sec_con_u']}. Appendices A.1 and A.2 contain supporting inequalities for Section \ref{['sec_con_b']} and \ref{['sec_con_u']}.

Theorems & Definitions (87)

Theorem 2.1.1
Remark 2.1.1
Theorem 2.2.1
Proposition 2.3.1
Theorem 3.1.1
Proposition 3.1.1
Lemma 3.2.1
proof
Lemma 3.2.2
proof
...and 77 more

Finite sample rates for logistic regression with small noise or few samples

TL;DR

Abstract

Finite sample rates for logistic regression with small noise or few samples

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (1)

Theorems & Definitions (87)