Fetching the paper…
Reading the bibliography…
Modern machine learning focuses on highly expressive models that are able to fit or interpolate the data completely, resulting in zero training loss.
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
On convergence proofs for perceptrons
Albert B Novikoff · 1963
Earlier work this paper cites.
Gradient methods for minimizing functionals
Boris Teodorovich Polyak · 1963
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
Robert E Schapire, Yoav Freund, Peter Bartlett, Wee Sun Lee, et al · 1998
Earlier work this paper cites.
Incremental gradient algorithms with stepsizes bounded away from zero
Mikhail V Solodov · 1998
Earlier work this paper cites.
An incremental gradient (-projection) method with momentum term and adaptive stepsize rule
Paul Tseng · 1998
Earlier work this paper cites.
Large-scale sparse logistic regression
Jun Liu, Jianhui Chen, and Jieping Ye · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Yu Nesterov · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Gradient methods for minimizing composite functions
Yu Nesterov · 2013
Cited alongside, same era.
Fast convergence of stochastic gradient descent under a strong growth condition
Mark Schmidt and Nicolas Le Roux · 2013
Cited alongside, same era.
A primal–dual smooth perceptron–von neumann algorithm
Negar Soheili and Javier Pena · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Later among the works it cites.
Convex until proven guilty: Dimension-free acceleration of gradient descent on non-convex functions
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2017
Later among the works it cites.
Accurate, large minibatch sgd: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Saddle points and accelerated perceptron algorithms
Adams Wei Yu, Fatma Kilinc-Karzan, and Jaime Carbonell · 2014
Cited alongside, same era.
Solving random quadratic systems of equations is nearly as easy as solving linear systems
Yuxin Chen and Emmanuel Candes · 2015
Cited alongside, same era.
Un-regularizing: approximate proximal point and faster stochastic algorithms for empirical risk minimization
Roy Frostig, Rong Ge, Sham Kakade, and Aaron Sidford · 2015
Cited alongside, same era.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2015
Cited alongside, same era.
Coordinate descent algorithms
Stephen J Wright · 2015
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Natasha 2: Faster non-convex optimization than sgd
Zeyuan Allen-Zhu · 2018
Closest in time.
On exponential convergence of sgd in non-convex over-parametrized learning
Raef Bassily, Mikhail Belkin, and Siyuan Ma · 2018
Closest in time.
On the linear convergence of the stochastic gradient method with constant step-size
Volkan Cevher and Bang Công Vũ · 2018
Closest in time.
On acceleration with noise-corrupted gradients
Michael Cohen, Jelena Diakonikolas, and Lorenzo Orecchia · 2018
Closest in time.
Accelerating stochastic gradient descent for least squares regression
Prateek Jain, Sham M. Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Closest in time.
An alternative view: When does sgd escape local minima?
Robert Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Closest in time.
Mass: an accelerated stochastic method for over-parametrized learning
Chaoyue Liu and Mikhail Belkin · 2018
Closest in time.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2018
Closest in time.