Fetching the paper…
Reading the bibliography…
We give a description of the high-dimensional limit of one-pass single-batch stochastic gradient descent (SGD) on a least squares problem.
“ A Stochastic Approximation Method ”
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
“ Random features for large-scale kernel machines ”
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
“Markov processes: characterization and convergence”
Stewart Ethier and Thomas Kurtz · 2009
Earlier work this paper cites.
“Toward a Noncommutative Arithmetic-geometric Mean Inequality: Conjectures, Case-studies, and Consequences”
Benjamin Recht and Christopher Re · 2012
Earlier work this paper cites.
“A note on the Hanson-Wright inequality for random vectors with dependencies”
R. Adamczak · 2015
Earlier work this paper cites.
Chuang Wang, Jonathan Mattingly and Yue. Lu · 2017
Earlier work this paper cites.
“High-dimensional probability: An introduction with applications in data science”
Roman Vershynin · 2018
Earlier work this paper cites.
“The scaling limit of high-dimensional online independent component analysis*”
Chuang Wang and Yue Lu · 2019
Cited alongside, same era.
“SGD with shuffling: optimal rates without component convexity and large epoch requirements”
Kwangjun Ahn, Chulhee Yun and Suvrit Sra · 2020
Cited alongside, same era.
“How Good is SGD with Random Shuffling?”
Itay Safran and Ohad Shamir · 2020
Cited alongside, same era.
“Online stochastic gradient descent on non-convex losses from high-dimensional inference”
Gerard Arous, Reza Gheissari and Aukosh Jagannath · 2021
Cited alongside, same era.
“The high-dimensional asymptotics of first order methods with random data”
Michael Celentano, Chen Cheng and Andrea Montanari · 2021
Cited alongside, same era.
“Why random reshuffling beats stochastic gradient descent”
“Open problem: Can single-shuffle SGD be better than reshuffling SGD and gd?”
Chulhee Yun, Suvrit Sra and Ali Jadbabaie · 2021
Later among the works it cites.
“High-dimensional limit theorems for SGD: Effective dynamics and critical scaling”
Gerard Arous, Reza Gheissari and Aukosh Jagannath · 2022
Later among the works it cites.
“Rigorous dynamical mean field theory for stochastic gradient descent methods”
Cedric Gerbelot, Emanuele Troiani, Francesca Mignacco, Florent Krzakala and Lenka Zdeborova · 2022
Later among the works it cites.
“Trajectory of Mini-Batch Momentum: Batch Size Saturation and Convergence in High Dimensions”
Kiwon Lee, Andrew. Cheng, Courtney Paquette and Elliot Paquette · 2022
Later among the works it cites.
“Homogenization of SGD in high-dimensions: Exact dynamics and generalization properties”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mert Gürbüzbalaban, Asu Ozdaglar and Pablo Parrilo · 2021
Cited alongside, same era.
“SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality”
Courtney Paquette, Kiwon Lee, Fabian Pedregosa and Elliot Paquette · 2021
Cited alongside, same era.
Courtney Paquette, Elliot Paquette, Ben Adlam and Jeffrey Pennington · 2022
Later among the works it cites.
Krishnakumar Balasubramanian, Promit Ghosal and Ye He · 2023
Closest in time.