Fetching the paper…
Reading the bibliography…
We develop a stochastic differential equation, called homogenized SGD, for analyzing the dynamics of stochastic gradient descent (SGD) on a high-dimensional random least squares problem with $\ell^2$-regularization.
A Stochastic Approximation Method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
L. Ljung · 1977
Earlier work this paper cites.
On the resolvents of nonconvolution Volterra kernels
G. Gripenberg · 1980
Earlier work this paper cites.
Anomalous diffusion in disordered media: statistical mechanisms, models and physical applications
J.-P. Bouchaud and A. Georges · 1990
Earlier work this paper cites.
Stochastic approximation and optimization of random systems , volume 17 of DMV Seminar
L. Ljung, G. Pflug, and H. Walk · 1992
Earlier work this paper cites.
Priors for Infinite Networks , pages 29–53
R.M. Neal · 1996
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Continuous martingales and Brownian motion , volume 293 of Grundlehren der mathematischen Wissenschaften [Fundamental Principles of Mathematical Sciences]
D. Revuz and M. Yor · 1999
Earlier work this paper cites.
Applied probability and queues , volume 51 of Applications of Mathematics (New York)
S. Asmussen · 2003
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
H. Kushner and G.G. Yin · 2003
Earlier work this paper cites.
Stochastic integration and differential equations , volume 21 of Stochastic Modelling and Applied Probability
P.E. Protter · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
Spectral analysis of large dimensional random matrices , volume 20
Z. Bai and J. Silverstein · 2010
Earlier work this paper cites.
Non-Asymptotic Analysis of Stochastic Approximation Algorithms for Machine Learning
E. Moulines and F. Bach · 2011
Earlier work this paper cites.
Random design analysis of ridge regression
D. Hsu, S. Kakade, and T. Zhang · 2012
Earlier work this paper cites.
In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
B. Neyshabur, R/ Tomioka, and N. Srebro · 2014
Earlier work this paper cites.
A note on the Hanson-Wright inequality for random vectors with dependencies
R. Adamczak · 2015
Earlier work this paper cites.
Averaged Least-Mean-Squares: Bias-Variance Trade-offs and Optimal Sampling Distributions
A. Defossez and F. Bach · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Earlier work this paper cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
N. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. Tang · 2016
Earlier work this paper cites.
Tracy-Widom distribution for the largest eigenvalue of real sample covariance matrices with general population
J. Lee and K. Schnelli · 2016
Earlier work this paper cites.
A variational analysis of stochastic gradient algorithms
S. Mandt, M. Hoffman, and D. Blei · 2016
Earlier work this paper cites.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
D. Needell, N. Srebro, and R. Ward · 2016
Earlier work this paper cites.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2016
Earlier work this paper cites.
Harder, Better, Faster, Stronger Convergence Rates for Least-Squares Regression
A. Dieuleveut, N. Flammarion, and F. Bach · 2017
Earlier work this paper cites.
A dynamical approach to random matrix theory , volume 28 of Courant Lecture Notes in Mathematics
L. Erdős and H-T. Yau · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
E. Hoffer, I. Hubara, and D. Soudry · 2017
Cited alongside, same era.
Three Factors Influencing Minima in SGD
S. Jastrzebski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2017
Cited alongside, same era.
Stochastic Modified Equations and Adaptive Stochastic Gradient Algorithms
Q. Li, C. Tai, and W. E · 2017
Cited alongside, same era.
Optimal Rates for Multi-pass Stochastic Gradient Methods
J. Lin and L. Rosasco · 2017
Cited alongside, same era.
Nonlinear random matrix theory for deep learning
J. Pennington and P. Worah · 2017
Cited alongside, same era.
Sharp convergence rates for Langevin dynamics in the nonconvex setting
The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization
D. Kobak, J. Lomond, and B. Sanchez · 2020
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
A. Lewkowycz, Y. Bahri, E. Dyer, J. Sohl-Dickstein, and G. Gur-Ari · 2020
Later among the works it cites.
Z. Liao, R. Couillet, and M. Mahoney · 2020
Later among the works it cites.
Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification
F. Mignacco, F. Krzakala, P. Urbani, and L. Zdeborová · 2020
Later among the works it cites.
Neural Kernels Without Tangents
V. Shankar, A. Fang, W. Guo, S. Fridovich-Keil, J. Ragan-Kelley, L. Schmidt, and B. Recht · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Cheng, N. Chatterji, Y. Abbasi-Yadkori, P. Bartlett, and M. Jordan · 2018
Cited alongside, same era.
High-dimensional asymptotics of prediction: ridge regression and classification
E. Dobriban and S. Wager · 2018
Cited alongside, same era.
Characterizing Implicit Bias in Terms of Optimization Geometry
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Accelerating Stochastic Gradient Descent for Least Squares Regression
P. Jain, S. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2018
Cited alongside, same era.
Deep Neural Networks as Gaussian Processes
J. Lee, Y. Bahri, R. Novak, S. Schoenholz, J. Pennington, and J. Sohl-Dickstein · 2018
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes
C. Li, H. Farkhoor, R. Liu, and J. Yosinski · 2018
Cited alongside, same era.
Later among the works it cites.
On the Generalization Benefit of Noise in Stochastic Gradient Descent
S. Smith, E. Elsen, and S. De · 2020
Later among the works it cites.
Benign overfitting in ridge regression
A. Tsigler and P. L. Bartlett · 2020
Later among the works it cites.
On the Optimal Weighted \ell_2 Regularization in Overparameterized Linear Regression
D. Wu and J. Xu · 2020
Later among the works it cites.
Implicit Gradient Regularization
D. Barrett and B. Dherin · 2021
Later among the works it cites.
Rank-one matrix estimation: analytic time evolution of gradient descent dynamics
A. Bodin and N. Macris · 2021
Later among the works it cites.
The high-dimensional asymptotics of first order methods with random data
M. Celentano, C. Cheng, and A. Montanari · 2021
Later among the works it cites.
Generalization Performance of Multi-pass Stochastic Gradient Descent with Convex Loss Functions
Y. Lei, T. Hu, and K. Tang · 2021
Later among the works it cites.
Logarithmic landscape and power-law escape rate of sgd
T. Mori, L. Ziyin, K. Liu, and M. Ueda · 2021
Later among the works it cites.
SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality
C. Paquette, K. Lee, F. Pedregosa, and E. Paquette · 2021
Later among the works it cites.
Asymptotics of Ridge(less) Regression under General Source Condition
D. Richards, J. Mourtada, and L. Rosasco · 2021
Later among the works it cites.
Lower bounds on the generalization error of nonlinear learning models
I. Seroussi and O. Zeitouni · 2021
Later among the works it cites.
Covariate Shift in High-Dimensional Random Feature Regression
N. Tripuraneni, B. Adlam, and J. Pennington · 2021
Later among the works it cites.
Last iterate convergence of SGD for Least-Squares in the Interpolation regime
A. Vardhan Varre, L. Pillaud-Vivien, and N. Flammarion · 2021
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Later among the works it cites.
The Benefits of Implicit Regularization from SGD in Least Squares Problems (NeurIPS)
D. Zou, J. Wu, V. Braverman, Q. Gu, D.P. Foster, and S. Kakade · 2021
Later among the works it cites.
Learning Curves for SGD on Structured Features
B. Bordelon and C. Pehlevan · 2022
Closest in time.
Random Matrix Methods for Machine Learning
R. Couillet and Z. Liao · 2022
Closest in time.
The generalization error of random features regression: precise asymptotics and the double descent curve
S. Mei and A. Montanari · 2022
Closest in time.
Strength of Minibatch Noise in SGD
L. Ziyin, K. Liu, T. Mori, and M. Ueda · 2022
Closest in time.
Risk Bounds of Multi-Pass SGD for Least Squares in the Interpolation Regime
D. Zou, J. Wu, V. Braverman, Q. Gu, and S. M. Kakade · 2022
Closest in time.