Fetching the paper…
Reading the bibliography…
Large-scale nonconvex optimization problems are ubiquitous in modern machine learning, and among practitioners interested in solving them, Stochastic Gradient Descent (SGD) reigns supreme.
Problem Complexity and Method Efficiency in Optimization
Arkadi Nemirovsky and David B. Yudin · 1983
Earlier work this paper cites.
Incremental Gradient Algorithms with Stepsizes Bounded Away from Zero
M.V. Solodov · 1998
Earlier work this paper cites.
An Incremental Gradient(-Projection) Method with Momentum Term and Adaptive Stepsize Rule
Paul Tseng · 1998
Earlier work this paper cites.
Gradient Convergence in Gradient methods with Errors
Dimitri P. Bertsekas and John N. Tsitsiklis · 2000
Earlier work this paper cites.
LibSVM: A library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
Optimal Distributed Online Prediction Using Mini-Batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Making Gradient Descent Optimal for Strongly Convex Stochastic Optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Earlier work this paper cites.
Stochastic First- and Zeroth-Order Methods for Nonconvex Stochastic Programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Fast Convergence of Stochastic Gradient Descent under a Strong Growth Condition
Mark Schmidt and Nicolas Le Roux · 2013
Earlier work this paper cites.
Stochastic Gradient Descent for Non-Smooth Optimization: Convergence Results and Optimal Averaging Schemes
Ohad Shamir and Tong Zhang · 2013
Earlier work this paper cites.
SAGA: A Fast Incremental Gradient Method with Support for Non-Strongly Convex Composite Objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Understanding machine learning: from theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Escaping From Saddle Points — Online Stochastic Gradient for Tensor Decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Randomized iterative methods for linear systems
Robert Mansel Gower and Peter Richtárik · 2015
Earlier work this paper cites.
Deep Learning with Limited Numerical Precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Stochastic Optimization with Importance Sampling for Regularized Loss Minimization
Peilin Zhao and Tong Zhang · 2015
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak- L
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
Deanna Needell, Nathan Srebro, and Rachel Ward · 2016
Cited alongside, same era.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
First-Order Methods in Optimization
Amir. Beck · 2017
Cited alongside, same era.
Communication-Efficient Learning of Deep Networks from Decentralized Data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
SGD: General Analysis and Improved Rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Later among the works it cites.
Convergence Rates for Deterministic and Stochastic Subgradient Methods without Lipschitz Continuity
Benjamin Grimmer · 2019
Later among the works it cites.
Stochastic distributed learning with gradient quantization and variance reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richtárik · 2019
Later among the works it cites.
Stochastic Gradient Descent for Nonconvex Learning Without Bounded Gradient Assumptions
Yunwei Lei, Ting Hu, Guiying Li, and Ke Tang · 2019
Later among the works it cites.
On the Convergence of Stochastic Gradient Descent with Adaptive Stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Batched Stochastic Gradient Descent with Weighted Sampling
Deanna Needell and Rachel Ward · 2017
Cited alongside, same era.
Stochastic Reformulations of Linear Systems: Algorithms and Convergence Theory
Peter Richtárik and Martin Takáč · 2017
Cited alongside, same era.
Optimization Methods for Large-Scale Machine Learning
Léon. Bottou, Frank E. Curtis, and Jorge. Nocedal · 2018
Cited alongside, same era.
Error Bounds, Quadratic Growth, and Linear Convergence of Proximal Methods
Dmitriy Drusvyatskiy and Adrian S. Lewis · 2018
Cited alongside, same era.
Distributed learning with compressed gradients
Sarit Khirirat, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2018
Cited alongside, same era.
Stochastic spectral and conjugate descent methods
Dmitry Kovalev, Eduard Gorbunov, Elnur Gasanov, and Peter Richtárik · 2018
Cited alongside, same era.
Linear Convergence of First Order Methods for Non-Strongly Convex Optimization
I Necoara, Y. Nesterov, and F. Glineur · 2019
Later among the works it cites.
Tight Dimension Independent Lower Bound on the Expected Convergence Rate for Diminishing Step Sizes in SGD
Phuong Ha Nguyen, Lam Nguyen, and Marten van Dijk · 2019
Later among the works it cites.
Karthik A. Sankararaman, Soham De, Zheng Xu, W. Ronny Huang, and Tom Goldstein · 2019
Later among the works it cites.
Unified Optimal Analysis of the (Stochastic) Gradient Method
Sebastian U. Stich · 2019
Later among the works it cites.
Sebastian U. Stich and Sai Praneeth Karimireddy · 2019
Later among the works it cites.
Hybrid Stochastic Gradient Descent Algorithms for Stochastic Nonconvex Optimization
Quoc Tran-Dinh, Nhan H. Pham, Dzung T. Phan, and Lam M. Nguyen · 2019
Later among the works it cites.
Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Later among the works it cites.
A Unified Theory of SGD: Variance Reduction, Sampling, Quantization and Coordinate Descent
Eduard Gorbunov, Filip Hanzely, and Peter Richtárik · 2020
Closest in time.
Stochastic quasi-gradient methods: variance reduction via Jacobian sketching
Robert M Gower, Peter Richtárik, and Francis Bach · 2020
Closest in time.
Optimization for Deep Learning: An Overview
Ruo-Yu Sun · 2020
Closest in time.
Complexity Issues in Global Optimization: A Survey
Stephen A. Vavasis · 2025
Closest in time.