Fetching the paper…
Reading the bibliography…
We analyze the dynamics of streaming stochastic gradient descent (SGD) in the high-dimensional limit when applied to generalized linear models and multi-index models (e.g.
Distribution of eigenvalues for some sets of random matrices
V.A. Marčenko and L.A. Pastur · 1967
Earlier work this paper cites.
Markov processes – characterization and convergence
Stewart N. Ethier and Thomas G. Kurtz · 1986
Earlier work this paper cites.
On milman’s inequality and random subspaces which escape through a mesh in ℝ n \mathbb{R}^{n}
Yehoram Gordon · 1988
Earlier work this paper cites.
On-line learning with a perceptron
Michael Biehl and Peter Riegler · 1994
Earlier work this paper cites.
Learning by on-line gradient descent
Michael Biehl and Holm Schwarze · 1995
Earlier work this paper cites.
Stochastic integration and differential equations , volume 21 of Stochastic Modelling and Applied Probability
P.E. Protter · 2005
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas Le Roux, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
M. Schmidt and N. Le Roux · 2013
Earlier work this paper cites.
Phase retrieval via wirtinger flow: Theory and algorithms
Emmanuel J Candes, Xiaodong Li, and Mahdi Soltanolkotabi · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
First-order methods in optimization
Amir Beck · 2017
Earlier work this paper cites.
Harder, better, faster, stronger convergence rates for least-squares regression
A. Dieuleveut, N. Flammarion, and F. Bach · 2017
Earlier work this paper cites.
Chuang Wang, Jonathan Mattingly, and Yue M Lu · 2017
Earlier work this paper cites.
Optimization methods for large-scale machine learning
L. Bottou, F.E. Curtis, and J. Nocedal · 2018
Earlier work this paper cites.
The Dynamics of Learning: A Random Matrix Approach
Z. Liao and R. Couillet · 2018
Earlier work this paper cites.
A random matrix approach to neural networks
C. Louart, Z. Liao, and R. Couillet · 2018
Cited alongside, same era.
Exponential convergence of testing error for stochastic gradient methods
Loucas Pillaud-Vivien, Alessandro Rudi, and Francis Bach · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science
R. Vershynin · 2018
Cited alongside, same era.
Optimal errors and phase transitions in high-dimensional generalized linear models
Jean Barbier, Florent Krzakala, Nicolas Macris, Léo Miolane, and Lenka Zdeborová · 2019
Cited alongside, same era.
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
Sebastian Goldt, Madhu Advani, Andrew M Saxe, Florent Krzakala, and Lenka Zdeborová · 2019
Cited alongside, same era.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Later among the works it cites.
The high-dimensional asymptotics of first order methods with random data, 2021
Michael Celentano, Chen Cheng, and Andrea Montanari · 2021
Later among the works it cites.
A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent
Z. Liao, R. Couillet, and M. Mahoney · 2021
Later among the works it cites.
Approximate message passing with spectral initialization for generalized linear models
Marco Mondelli and Ramji Venkataramanan · 2021
Later among the works it cites.
Dynamics of stochastic momentum methods on large-scale, quadratic models
Courtney Paquette and Elliot Paquette · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan · 2019
Cited alongside, same era.
The impact of regularization on high-dimensional logistic regression
Fariborz Salehi, Ehsan Abbasi, and Babak Hassibi · 2019
Cited alongside, same era.
A solvable high-dimensional model of GAN
Chuang Wang, Hong Hu, and Yue Lu · 2019
Cited alongside, same era.
Yuki Yoshida and Masato Okada · 2019
Cited alongside, same era.
The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression
Emmanuel J Candès and Pragya Sur · 2020
Cited alongside, same era.
The nonsmooth landscape of phase retrieval
Damek Davis, Dmitriy Drusvyatskiy, and Courtney Paquette · 2020
Cited alongside, same era.
Modeling the influence of data structure on learning in neural networks: The hidden manifold model
Sebastian Goldt, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2020
Cited alongside, same era.
Atish Agarwala, Fabian Pedregosa, and Jeffrey Pennington · 2022
Later among the works it cites.
High-dimensional limit theorems for sgd: Effective dynamics and critical scaling
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2022
Later among the works it cites.
Learning single-index models with shallow neural networks
Alberto Bietti, Joan Bruna, Clayton Sanford, and Min Jae Song · 2022
Later among the works it cites.
Rigorous dynamical mean field theory for stochastic gradient descent methods, 2022
Cedric Gerbelot, Emanuele Troiani, Francesca Mignacco, Florent Krzakala, and Lenka Zdeborova · 2022
Later among the works it cites.
The gaussian equivalence of generative models for learning with shallow neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2022
Later among the works it cites.
Implicit Regularization or Implicit Conditioning? Exact Risk Trajectories of SGD in High Dimensions
Courtney Paquette, Elliot Paquette, Ben Adlam, and Jeffrey Pennington · 2022
Later among the works it cites.
Sharp global convergence guarantees for iterative nonconvex optimization with random data
Kabir Aladin Chandrasekher, Ashwin Pananjady, and Christos Thrampoulidis · 2023
Closest in time.
High-dimensional limit of one-pass SGD on least squares
Elizabeth Collins-Woodfin and Elliot Paquette · 2023
Closest in time.
Smoothing the landscape boosts the signal for SGD: Optimal sample complexity for learning single index models, 2023
Alex Damian, Eshaan Nichani, Rong Ge, and Jason D. Lee · 2023
Closest in time.
Neural networks efficiently learn low-dimensional representations with SGD, 2023
Alireza Mousavi-Hosseini, Sejun Park, Manuela Girotti, Ioannis Mitliagkas, and Murat A. Erdogdu · 2023
Closest in time.
Online stochastic gradient descent with arbitrary initialization solves non-smooth, non-convex phase retrieval
Yan Shuo Tan and Roman Vershynin · 2023
Closest in time.