Fetching the paper…
Reading the bibliography…
We propose a new framework, inspired by random matrix theory, for analyzing the dynamics of stochastic gradient descent (SGD) when both number of samples and dimensions are large.
A Stochastic Approximation Method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Quicksort
C.A.R. Hoare · 1962
Earlier work this paper cites.
Distribution of eigenvalues for some sets of random matrices
V.A. Marčenko and L.A. Pastur · 1967
Earlier work this paper cites.
On optional stochastic integrals and a remarkable series of exponential formulas
M. Yor · 1976
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
L. Ljung · 1977
Earlier work this paper cites.
Sur le comportement asymptotique des martingales locales
D. Lépingle · 1978
Earlier work this paper cites.
Random-energy model: An exactly solvable model of disordered systems
B Derrida · 1981
Earlier work this paper cites.
On the average number of steps of the simplex method of linear programming
S. Smale · 1983
Earlier work this paper cites.
A Probabilistic Analysis of the Simplex Method
K. Borgwardt · 1986
Earlier work this paper cites.
Stochastic minimization with constant step-size: asymptotic laws
G. C. Pflug · 1986
Earlier work this paper cites.
Empirical processes with applications to statistics
G. Shorack and J. Wellner · 1986
Earlier work this paper cites.
Volterra Integral and Functional Equations
G. Gripenberg, S.O. Londen, and O. Staffans · 1990
Earlier work this paper cites.
Adventures in stochastic processes
S. Resnick · 1992
Earlier work this paper cites.
Poisson Processes
J. F. C. Kingman · 1993
Earlier work this paper cites.
Information-based complexity of convex programming
A. Nemirovski · 1995
Earlier work this paper cites.
A new class of incremental gradient methods for least squares problems
D. Bertsekas · 1997
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
D. Bertsekas and J. Tsitsiklis · 2000
Earlier work this paper cites.
Bounds on Tail Probabilities of Discrete Distributions
B. Klar · 2000
Earlier work this paper cites.
Applied probability and queues , volume 51 of Applications of Mathematics (New York)
S. Asmussen · 2003
Earlier work this paper cites.
Asymptotics for sums of random variables with local subexponential behaviour
S. Asmussen, S. Foss, and D. Korshunov · 2003
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
H. Kushner and G.G. Yin · 2003
Earlier work this paper cites.
Introductory lectures on convex optimization
Y. Nesterov · 2004
Earlier work this paper cites.
Smoothed Analysis of Algorithms: Why the Simplex Algorithm Usually Takes Polynomial Time
D. Spielman and S. Teng · 2004
Earlier work this paper cites.
Stochastic integration and differential equations , volume 21 of Stochastic Modelling and Applied Probability
P.E. Protter · 2005
Earlier work this paper cites.
Freezing and extreme-value statistics in a random energy model with logarithmically correlated potential
Y. Fyodorov and J. Bouchaud · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
A randomized Kaczmarz algorithm with exponential convergence
T. Strohmer and R. Vershynin · 2009
Earlier work this paper cites.
Asymptotic distribution of singular values of powers of random matrices
N. Alexeev, F. Götze, and A. Tikhomirov · 2010
Cited alongside, same era.
Spectral analysis of large dimensional random matrices , volume 20
Z. Bai and J. Silverstein · 2010
Cited alongside, same era.
On explicit probability densities associated with Fuss-Catalan numbers
D. Liu, C. Song, and Z. Wang · 2011
Cited alongside, same era.
Hybrid deterministic-stochastic methods for data fitting
M. Friedlander and M. Schmidt · 2012
Cited alongside, same era.
Topics in random matrix theory , volume 132
T. Tao · 2012
Cited alongside, same era.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P. Nguyen · 2018
Later among the works it cites.
The full spectrum of deepnet hessians at scale: Dynamics with SGD Training and Sample Size
V. Papyan · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science
R. Vershynin · 2018
Later among the works it cites.
An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
G. Behrooz, S. Krishnan, and Y. Xiao · 2019
Later among the works it cites.
Eigenvalue distribution of nonlinear models of random matrices
L. Benigni and S. Péché · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. Saxe, J. McClelland, and S. Ganguli · 2013
Cited alongside, same era.
No more pesky learning rates
T. Schaul, S. Zhang, and Y. LeCun · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y.N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
N. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. Tang · 2016
Cited alongside, same era.
A variational analysis of stochastic gradient algorithms
S. Mandt, M. Hoffman, and D. Blei · 2016
Cited alongside, same era.
R. Gower, N. Loizou, X. Qian, A. Sailanbayev, E. Shulgin, and P. Richtárik · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
T. Hastie, A. Montanari, S. Rosset, and R.J. Tibshirani · 2019
Later among the works it cites.
The random matrix theory of the classical compact groups , volume 218 of Cambridge Tracts in Mathematics
E. Meckes · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
S. Mei and A. Montanari · 2019
Later among the works it cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Later among the works it cites.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise
T. Nguyen, U. Simsekli, M. Gurbuzbalaban, and G. Richard · 2019
Later among the works it cites.
A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks
U. Simsekli, L. Sagun, and M. Gurbuzbalaban · 2019
Later among the works it cites.
Painless Stochastic Gradient: Interpolation, Line-Search, and Convergence Rates
S. Vaswani, A. Mishkin, I. Laradji, M. Schmidt, G. Gidel, and S. Lacoste-Julien · 2019
Later among the works it cites.
Fluctuation-dissipation relations for stochastic gradient descent
S. Yaida · 2019
Later among the works it cites.
Z. Zhu, J. Wu, B. Yu, L. Wu, and J. Ma · 2019
Later among the works it cites.
Eigenstate Thermalization Hypothesis for Wigner Matrices
G. Cipolloni, L. Erdős, and D. Schröder · 2020
Later among the works it cites.
The conjugate gradient algorithm on well-conditioned Wishart matrices is almost deteriministic
P.A. Deift and T. Trogdon · 2020
Later among the works it cites.
Asymptotic Errors for High-Dimensional Convex Penalized Linear Regression beyond Gaussian Matrices
C. Gerbelot, A. Abbara, and F. Krzakala · 2020
Later among the works it cites.
The Heavy-Tail Phenomenon in SGD
M. Gurbuzbalaban, U. Simsekli, and L. Zhu · 2020
Later among the works it cites.
Matrix Concentration for Products
D. Huang, J. Niles-Weed, J. Tropp, and R. Ward · 2020
Later among the works it cites.
Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics
D. Kunin, J. Sagastuy-Brena, S. Ganguli, D.L. K. Yamins, and H. Tanaka · 2020
Later among the works it cites.
Optimal Randomized First-Order Methods for Least-Squares Problems
J. Lacotte and M. Pilanci · 2020
Later among the works it cites.
Halting Time is Predictable for Large Models: A Universality Property and Average-case Analysis
C. Paquette, B. van Merriënboer, and F. Pedregosa · 2020
Later among the works it cites.
Universality for the conjugate gradient and MINRES algorithms on sample covariance matrices
E. Paquette and T. Trogdon · 2020
Later among the works it cites.
Average-case Acceleration Through Spectral Density Estimation
F. Pedregosa and D. Scieur · 2020
Later among the works it cites.
Mean field analysis of neural networks: a law of large numbers
J. Sirignano and K. Spiliopoulos · 2020
Later among the works it cites.
Randomized Kaczmarz converges along small singular vectors
S. Steinerberger · 2020
Later among the works it cites.