Fetching the paper…
Reading the bibliography…
Random Reshuffling (RR) is an algorithm for minimizing finite-sum functions that utilizes iterative gradient descent steps in conjunction with data reshuffling.
On the Convergence of the LMS Algorithm with Adaptive Learning Rate for Linear Feedforward Networks
Zhi-Quan Luo · 1991
Earlier work this paper cites.
Convergence Analysis of a Proximal-Like Minimization Algorithm Using Bregman Functions
Gong Chen and Marc Teboulle · 1993
Earlier work this paper cites.
A class of unconstrained minimization methods for neural network training
Luigi Grippo · 1994
Earlier work this paper cites.
Serial and parallel backpropagation convergence via nonmonotone perturbed minimization
Olvi L. Mangasarian and Mikhail V. Solodov · 1994
Earlier work this paper cites.
Gradient Convergence in Gradient methods with Errors
Dimitri P. Bertsekas and John N. Tsitsiklis · 2000
Earlier work this paper cites.
Incremental Subgradient Methods for Nondifferentiable Optimization
Angelia Nedić and Dimitri P. Bertsekas · 2001
Earlier work this paper cites.
Curiously fast convergence of some stochastic gradient descent algorithms
Léon Bottou · 2009
Earlier work this paper cites.
Incremental Gradient, Subgradient, and Proximal Methods for Convex Optimization: A Survey
Dimitri P. Bertsekas · 2011
Earlier work this paper cites.
Practical Recommendations for Gradient-Based Training of Deep Architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
Making Gradient Descent Optimal for Strongly Convex Stochastic Optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Earlier work this paper cites.
Toward a noncommutative arithmetic-geometric mean inequality: Conjectures, case-studies, and consequences
Benjamin Recht and Christopher Ré · 2012
Earlier work this paper cites.
Parallel Stochastic Gradient Algorithms for Large-Scale Matrix Completion
Benjamin Recht and Christopher Ré · 2013
Cited alongside, same era.
Stochastic Gradient Descent, Weighted Sampling, and the Randomized Kaczmarz algorithm
Deanna Needell, Rachel Ward, and Nati Srebro · 2014
Cited alongside, same era.
Optimization Methods for Large-Scale Machine Learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
SGD and Hogwild! Convergence Without the Bounded Gradients Assumption
Lam Nguyen, Phuong Ha Nguyen, Marten van Dijk, Peter Richtárik, Katya Scheinberg, and Martin Takáč · 2018
Cited alongside, same era.
Stochastic Learning Under Random Reshuffling With Constant Step-Sizes
Bicheng Ying, Kun Yuan, Stefan Vlaski, and Ali H. Sayed · 2018
Cited alongside, same era.
The Complexity of Finding Stationary Points with Stochastic Gradient Descent
Tight Dimension Independent Lower Bound on the Expected Convergence Rate for Diminishing Step Sizes in SGD
Phuong Ha Nguyen, Lam Nguyen, and Marten van Dijk · 2019
Later among the works it cites.
Unified Optimal Analysis of the (Stochastic) Gradient Method
Sebastian U. Stich · 2019
Later among the works it cites.
On Tight Convergence Rates of Without-replacement SGD
Kwangjun Ahn and Suvrit Sra · 2020
Closest in time.
SGD with shuffling: optimal rates without component convexity and large epoch requirements
Kwangjun Ahn, Chulhee Yun, and Suvrit Sra · 2020
Closest in time.
Better Theory for SGD in the Nonconvex World
Ahmed Khaled and Peter Richtárik · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yoel Drori and Ohad Shamir · 2019
Cited alongside, same era.
SGD: General Analysis and Improved Rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Cited alongside, same era.
Random Shuffling Beats SGD after Finite Epochs
Jeff Haochen and Suvrit Sra · 2019
Cited alongside, same era.
Incremental Methods for Weakly Convex Optimization
Xiao Li, Zhihui Zhu, Anthony Man-Cho So, and Jason D Lee · 2019
Cited alongside, same era.
Adaptive gradient descent without descent
Yura Malitsky and Konstantin Mishchenko · 2019
Cited alongside, same era.
SGD without Replacement: Sharper Rates for General Smooth Convex Functions
Dheeraj Nagaraj, Prateek Jain, and Praneeth Netrapalli · 2019
Cited alongside, same era.
Convergence Rate of Incremental Gradient and Incremental Newton Methods
Mert Gürbüzbalaban, Asuman Ozdaglar, and Pablo A. Parrilo
Cited in the paper.
Closest in time.
Recht-Ré Noncommutative Arithmetic-Geometric Mean Conjecture is False
Zehua Lai and Lek-Heng Lim · 2020
Closest in time.
A Unified Convergence Analysis for Shuffling-Type Gradient Methods
Lam M. Nguyen, Quoc Tran-Dinh, Dzung T. Phan, Phuong Ha Nguyen, and Marten van Dijk · 2020
Closest in time.
Closing the convergence gap of SGD without replacement
Shashank Rajput, Anant Gupta, and Dimitris Papailiopoulos · 2020
Closest in time.
How Good is SGD with Random Shuffling?
Itay Safran and Ohad Shamir · 2020
Closest in time.
Optimization for Deep Learning: An Overview
Ruo-Yu Sun · 2020
Closest in time.