Accelerated Distance-adaptive Methods for Hölder Smooth and Convex Optimization
Yijin Ren, Haifeng Xu, Qi Deng
TL;DR
The paper tackles convex optimization where the objective has Hölder smoothness by introducing a parameter-free accelerated distance-adaptive method (AGDA) that estimates distance to the minimizer through distance-adaptive quantities and line-search guided updates, achieving anytime convergence without requiring $L_\nu$, $\nu$, or $D_0$. It further extends to a line-search-free stochastic variant (LF-AGDA) under a bounded-domain assumption, with convergence guarantees that mix local Hölder smoothness and stochastic noise. The proposed methods outperform existing universal or parameter-tuned schemes in deterministic and stochastic experiments, and exhibit robustness to mis-specified problem parameters, including the initial distance estimate $\bar{r}$. Overall, the work advances practical, adaptive first-order optimization for Hölder-smooth convex problems and provides empirically competitive results in nonconvex neural-network settings.
Abstract
This paper introduces new parameter-free first-order methods for convex optimization problems in which the objective function exhibits Hölder smoothness. Inspired by the recently proposed distance-over-gradient (DOG) technique, we propose an accelerated distance-adaptive method which achieves optimal anytime convergence rates for Hölder smooth problems without requiring prior knowledge of smoothness parameters or explicit parameter tuning. Importantly, our parameter-free approach removes the necessity of specifying target accuracy in advance, addressing a limitation found in the universal fast gradient methods (Nesterov, Yu. \textit{Mathematical Programming}, 2015). For convex stochastic optimization, we further present a parameter-free accelerated method that eliminates the need for line-search procedures. Preliminary experimental results highlight the effectiveness of our approach on convex nonsmooth problems and its advantages over existing parameter-free or accelerated methods.
