Fetching the paper…
Reading the bibliography…
We study the scaling limits of stochastic gradient descent (SGD) with constant step-size in the high-dimensional regime.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
An introduction to multivariate statistical analysis
Theodore Wilbur Anderson · 1962
Earlier work this paper cites.
Functional and random central limit theorems for the Robbins-Munro process
D. L. McLeish · 1976
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
Lennart Ljung · 1977
Earlier work this paper cites.
Asymptotic behavior of stochastic approximation and large deviations
Harold J Kushner · 1984
Earlier work this paper cites.
Markov processes
Stewart N. Ethier and Thomas G. Kurtz · 1986
Earlier work this paper cites.
Stochastic approximation and large deviations: Upper bounds and w.p.1 convergence
Paul Dupuis and Harold J Kushner · 1989
Earlier work this paper cites.
Adaptive algorithms and stochastic approximations
Albert Benveniste, Michel Métivier, and Pierre Priouret · 1990
Earlier work this paper cites.
The spherical p-spin interaction spin-glass model
A Crisanti, H Horner, and H-J Sommers · 1993
Earlier work this paper cites.
Analytical solution of the off-equilibrium dynamics of a long-range spin-glass model
Leticia F. Cugliandolo and Jorge Kurchan · 1993
Earlier work this paper cites.
Dynamics of on-line gradient descent learning for multilayer neural networks
David Saad and Sara Solla · 1995
Earlier work this paper cites.
On-line learning in soft committee machines
David Saad and Sara A Solla · 1995
Earlier work this paper cites.
Algorithmes stochastiques
Marie Duflo · 1996
Earlier work this paper cites.
Multidimensional diffusion processes
Daniel W. Stroock and S. R. Srinivasa Varadhan · 1997
Earlier work this paper cites.
Dynamics of stochastic approximation algorithms
Michel Benaïm · 1999
Earlier work this paper cites.
On-Line Learning and Stochastic Approximations
Léon Bottou · 1999
Earlier work this paper cites.
On the distribution of the largest eigenvalue in principal components analysis
Iain M Johnstone · 2001
Earlier work this paper cites.
Large scale online learning
Léon Bottou and Yan Le Cun · 2004
Earlier work this paper cites.
Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices
Jinho Baik, Gérard Ben Arous, and Sandrine Péché · 2005
Earlier work this paper cites.
The largest eigenvalue of rank one deformation of large wigner matrices
Delphine Féral and Sandrine Péché · 2007
Earlier work this paper cites.
Asymptotics of sample eigenstructure for a large dimensional spiked covariance model
Debashis Paul · 2007
Earlier work this paper cites.
The largest eigenvalues of finite rank deformation of large wigner matrices: convergence and nonuniversality of the fluctuations
Mireille Capitaine, Catherine Donati-Martin, and Delphine Féral · 2009
Earlier work this paper cites.
The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices
Florent Benaych-Georges and Raj Rao Nadakuditi · 2011
Earlier work this paper cites.
Ordinary differential equations and dynamical systems
Gerald Teschl · 2012
Earlier work this paper cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Deanna Needell, Nathan Srebro, and Rachel Ward · 2014
Earlier work this paper cites.
A statistical model for tensor PCA
Emile Richard and Andrea Montanari · 2014
Earlier work this paper cites.
Escaping from saddle points — online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Tensor principal component analysis via sum-of-square proofs
Samuel B Hopkins, Jonathan Shi, and David Steurer · 2015
Cited alongside, same era.
On the limitation of spectral methods: From the gaussian hidden clique problem to rank-one perturbations of gaussian tensors
Andrea Montanari, Daniel Reichman, and Ofer Zeitouni · 2015
Cited alongside, same era.
Fast spectral algorithms from sum-of-squares proofs: tensor decomposition and planted sparse vectors
Samuel B Hopkins, Tselil Schramm, Jonathan Shi, and David Steurer · 2016
Cited alongside, same era.
Online ICA: Understanding global dynamics of nonconvex optimization via diffusion processes
Chris Junchi Li, Zhaoran Wang, and Han Liu · 2016
Cited alongside, same era.
Tight analyses for non-smooth stochastic gradient descent
Nicholas J. A. Harvey, Christopher Liaw, Yaniv Plan, and Sikander Randhawa · 2019
Later among the works it cites.
Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations
Qianxiao Li, Cheng Tai, and E Weinan · 2019
Later among the works it cites.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
Vardan Papyan · 2019
Later among the works it cites.
Who is afraid of big bad minima? analysis of gradient-flow in spiked matrix-tensor models
Stefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota, Florent Krzakala, and Lenka Zdeborová · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint
Martin J Wainwright · 2019
Later among the works it cites.
Algorithmic thresholds for tensor PCA
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convergence of stochastic gradient descent for PCA
Ohad Shamir · 2016
Cited alongside, same era.
Community detection in hypergraphs, spiked tensor models, and sum-of-squares
Chiheon Kim, Afonso S Bandeira, and Michel X Goemans · 2017
Cited alongside, same era.
Statistical and computational phase transitions in spiked tensor estimation
Thibault Lesieur, Léo Miolane, Marc Lelarge, Florent Krzakala, and Lenka Zdeborová · 2017
Cited alongside, same era.
Diffusion approximations for online principal component estimation and global convergence
Chris Junchi Li, Mengdi Wang, Han Liu, and Tong Zhang · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate Bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Cited alongside, same era.
Perceptrons, Reissue of the 1988 Expanded Edition with a new foreword by Léon Bottou: An Introduction to Computational Geometry
Marvin Minsky and Seymour A Papert · 2017
Cited alongside, same era.
Non-convex learning via stochastic gradient Langevin dynamics: a nonasymptotic analysis
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky · 2017
Cited alongside, same era.
Gérard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2020
Later among the works it cites.
Bounding flows for spherical spin glass dynamics
Gérard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2020
Later among the works it cites.
Stochastic gradient and Langevin processes
Xiang Cheng, Dong Yin, Peter Bartlett, and Michael Jordan · 2020
Later among the works it cites.
Bridging the gap between constant step size stochastic gradient descent and Markov chains
Aymeric Dieuleveut, Alain Durmus, and Francis Bach · 2020
Later among the works it cites.
Fundamental limits of detection in the spiked Wigner model
Ahmed El Alaoui, Florent Krzakala, and Michael Jordan · 2020
Later among the works it cites.
Statistical thresholds for tensor PCA
Aukosh Jagannath, Patrick Lopatto, and Léo Miolane · 2020
Later among the works it cites.
Marvels and pitfalls of the Langevin algorithm in noisy high-dimensional inference
Stefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborová · 2020
Later among the works it cites.
Statistical limits of spiked tensor models
Amelia Perry, Alexander S. Wein, and Afonso S. Bandeira · 2020
Later among the works it cites.
Mean field analysis of neural networks: A central limit theorem
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Later among the works it cites.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Later among the works it cites.
The high-dimensional asymptotics of first order methods with random data
Michael Celentano, Chen Cheng, and Andrea Montanari · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
On the validity of modeling SGD with stochastic differential equations (SDEs)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora · 2021
Later among the works it cites.
SGD in the large: Average-case analysis, asymptotics, and stepsize criticality
Courtney Paquette, Kiwon Lee, Fabian Pedregosa, and Elliot Paquette · 2021
Later among the works it cites.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová · 2021
Later among the works it cites.
Learning threshold neurons via the "edge of stability", 2022
Kwangjun Ahn, Sebastien Bubeck, Sinho Chewi, Yin Tat Lee, Felipe Suarez, and Yi Zhang · 2022
Closest in time.
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Closest in time.
High-dimensional asymptotics of Langevin dynamics in spiked matrix models
Tengyuan Liang, Subhabrata Sen, and Pragya Sur · 2022
Closest in time.
Phase diagram of stochastic gradient descent in high-dimensional two-layer neural networks
Rodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2022
Closest in time.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D. Lee · 2023
Closest in time.
Understanding edge-of-stability training dynamics with a minimalist example
Xingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou, and Rong Ge · 2023
Closest in time.