Cheik Traoré

OC
h-index2
3papers
42citations
Novelty52%
AI Score43

3 Papers

8.9OCAug 18, 2023Code
Variance reduction techniques for stochastic proximal point algorithms

Cheik Traoré, Vassilis Apidopoulos, Saverio Salzo et al.

In the context of finite sums minimization, variance reduction techniques are widely used to improve the performance of state-of-the-art stochastic gradient methods. Their practical impact is clear, as well as their theoretical properties. Stochastic proximal point algorithms have been studied as an alternative to stochastic gradient algorithms since they are more stable with respect to the choice of the step size. However, their variance-reduced versions are not as well studied as the gradient ones. In this work, we propose the first unified study of variance reduction techniques for stochastic proximal point algorithms. We introduce a generic stochastic proximal-based algorithm that can be specified to give the proximal version of SVRG, SAGA, and some of their variants. For this algorithm, in the smooth setting, we provide several convergence rates for the iterates and the objective function values, which are faster than those of the vanilla stochastic proximal point algorithm. More specifically, for convex functions, we prove a sublinear convergence rate of $O(1/k)$. In addition, under the Polyak-Łojasiewicz (PL) condition, we obtain linear convergence rates. Finally, our numerical experiments demonstrate the advantages of the proximal variance reduction methods over their gradient counterparts in terms of the stability with respect to the choice of the step size in most cases, especially for difficult problems.

5.1OCMar 27
The adjoint state method for parametric definable optimization without smoothness or uniqueness

Jérôme Bolte, Edouard Pauwels, Cheik Traoré

Definable parametric optimization problems with possibly nonsmooth objectives, inequality constraints, and non-unique primal and dual solutions admit an adjoint state formula under a mere qualification condition. The adjoint construction yields a selection of a conservative field for the value function, providing a computable first-order object without requiring differentiation of the solution mapping. Through examples, we show that even in smooth problems, the formal adjoint construction fails without conservativity or definability, illustrating the relevance of these concepts to grasp theoretical aspects of the method. This work provides a tool which can be directly combined with existing primal-dual solvers for a wide range of parametric optimization problems.

8.0OCNov 24, 2020
Sequential convergence of AdaGrad algorithm for smooth convex optimization

Cheik Traoré, Edouard Pauwels

We prove that the iterates produced by, either the scalar step size variant, or the coordinatewise variant of AdaGrad algorithm, are convergent sequences when applied to convex objective functions with Lipschitz gradient. The key insight is to remark that such AdaGrad sequences satisfy a variable metric quasi-Fejér monotonicity property, which allows to prove convergence.