Fetching the paper…
Reading the bibliography…
The training of machine learning models is typically carried out using some form of gradient descent, often with great success.
Variable metric method for minimization
W. C. Davidon · 1959
Earlier work this paper cites.
Nonlinear functional analysis
K. Deimling · 1985
Earlier work this paper cites.
Matrix Analysis
Roger A. Horn and Charles R. Johnson, editors · 1986
Earlier work this paper cites.
Variable metric method for minimization
William C Davidon · 1991
Earlier work this paper cites.
Probability with Martingales
D. Williams · 1991
Earlier work this paper cites.
Iterative weighted least squares algorithms for neural networks classifiers
Takio Kurita · 1993
Earlier work this paper cites.
Sharp uniform convexity and smoothness inequalities for trace norms
E.H. Lieb, Keith Ball, and E.A. Carlen · 1994
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
Learning a nonlinear embedding by preserving class neighbourhood structure
Ruslan Salakhutdinov and Geoffrey E. Hinton · 2007
Earlier work this paper cites.
Optimization algorithms on matrix manifolds
P-A Absil, Robert Mahony, and Rodolphe Sepulchre · 2009
Earlier work this paper cites.
Geometric properties of Banach spaces and nonlinear iterations , volume 1965
Charles Chidume · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations
Honglak Lee, Roger Grosse, Rajesh Ranganath, and Andrew Y Ng · 2009
Earlier work this paper cites.
Adaptive cubic regularisation methods for unconstrained optimization. part i: motivation, convergence and numerical results
Coralia Cartis, Nicholas IM Gould, and Philippe L Toint · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R. Bach · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Manifolds, tensor analysis, and applications , volume 75
Ralph Abraham, Jerrold E Marsden, and Tudor Ratiu · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoff Hinton · 2012
Cited alongside, same era.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Learning hierarchical features for scene labeling
Clement Farabet, Camille Couprie, Laurent Najman, and Yann LeCun · 2013
Cited alongside, same era.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
A trust region algorithm with a worst-case iteration complexity of o ( ϵ − 3 / 2 ) ) o(\epsilon^{-3/2})) for nonconvex optimization
Frank E Curtis, Daniel P Robinson, and Mohammadreza Samadi · 2017
Closest in time.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Closest in time.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Closest in time.
First order methods beyond convexity and lipschitz gradient continuity with applications to quadratic inverse problems
Jérôme Bolte, Shoham Sabach, Marc Teboulle, and Yakov Vaisbourd · 2018
Closest in time.
Global rates of convergence for nonconvex optimization on manifolds
N. Boumal, P.-A. Absil, and C. Cartis · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Cited alongside, same era.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Multimodal learning with deep boltzmann machines
Nitish Srivastava and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
Riemannian metrics for neural networks i: feedforward networks
Yann Ollivier · 2015
Cited alongside, same era.
Stochastic model-based minimization under high-order growth
Damek Davis, Dmitriy Drusvyatskiy, and Kellie J MacPhee · 2018
Closest in time.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2018
Closest in time.
Relatively smooth convex optimization by first-order methods, and applications
H. Lu, R. Freund, and Y. Nesterov · 2018
Closest in time.
Sgd and hogwild: Convergence without the bounded gradients assumption
Lam et. al Nguyen · 2018
Closest in time.
On the convergence rate of stochastic mirror descent for nonsmooth nonconvex optimization
Siqi Zhang and Niao He · 2018
Closest in time.
On linear convergence of non-euclidean gradient methods without strong convexity and lipschitz gradient continuity
Heinz H. Bauschke, Jérôme Bolte, Jiawei Chen, Marc Teboulle, and Xianfu Wang · 2019
Closest in time.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2019
Closest in time.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Closest in time.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Benjamin Fehrman, Benjamin Gess, and Arnulf Jentzen · 2020
Closest in time.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2020
Closest in time.
Stability and convergence of stochastic gradient clipping: Beyond lipschitz continuity and smoothness
Vien V. Mai and Mikael Johansson · 2021
Closest in time.
On the convergence of mSGD and adagrad for stochastic optimization
Ruinan Jin, Yu Xing, and Xingkang He · 2022
Closest in time.
High probability bounds for a class of nonconvex algorithms with adagrad stepsize
Ali Kavis, Kfir Yehuda Levy, and Volkan Cevher · 2022
Closest in time.