Fetching the paper…
Reading the bibliography…
This manuscript investigates the one-pass stochastic gradient descent (SGD) dynamics of a two-layer neural network trained on Gaussian data and labels generated by a similar, though not necessarily identical, target function.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Stochastic minimization with constant step-size: Asymptotic laws
Georg Ch. Pflug · 1986
Earlier work this paper cites.
On-line learning in soft committee machines
David Saad and Sara A. Solla · 1995
Earlier work this paper cites.
Exact solution for on-line learning in multilayer neural networks
David Saad and Sara A. Solla · 1995
Earlier work this paper cites.
Learning by on-line gradient descent
M Biehl and H Schwarze · 1995
Earlier work this paper cites.
On-line learning in the committee machine
M Copelli and N Caticha · 1995
Earlier work this paper cites.
Dynamics of on-line gradient descent learning for multilayer neural networks
David Saad and Sara Solla · 1996
Earlier work this paper cites.
Transient dynamics of on-line learning in two-layered neural networks
Michael Biehl, Peter Riegler, and Christian Wöhler · 1996
Earlier work this paper cites.
Large scale online learning
Léon Bottou and Yann LeCun · 2003
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2007
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis Bach · 2011
Earlier work this paper cites.
Concentration Inequalities: A Nonasymptotic Theory of Independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
Harder, better, faster, stronger convergence rates for least-squares regression
Aymeric Dieuleveut, Nicolas Flammarion, and Francis Bach · 2017
Earlier work this paper cites.
A Practical Guide to Deterministic Particle Methods
A. Chertock · 2017
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaic Chizat and Francis Bach · 2018
Cited alongside, same era.
High-Dimensional Probability: An Introduction with Applications in Data Science
Roman Vershynin · 2018
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Cited alongside, same era.
Trainability and accuracy of neural networks: An interacting particle system approach, 2019
Benign overfitting of constant-stepsize sgd for linear regression
Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, and Sham Kakade · 2021
Later among the works it cites.
Last iterate convergence of sgd for least-squares in the interpolation regime
Aditya Vardhan Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Later among the works it cites.
The gaussian equivalence of generative models for learning with two-layer neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2021
Later among the works it cites.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborova · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grant M. Rotskoff and Eric Vanden-Eijnden · 2019
Cited alongside, same era.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
Sebastian Goldt, Madhu Advani, Andrew M Saxe, Florent Krzakala, and Lenka Zdeborová · 2019
Cited alongside, same era.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Cited alongside, same era.
Online stochastic gradient descent with arbitrary initialization solves non-smooth, non-convex phase retrieval, 2019
Yan Shuo Tan and Roman Vershynin · 2019
Cited alongside, same era.
Mean field analysis of neural networks: A central limit theorem
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Cited alongside, same era.
Bridging the gap between constant step size stochastic gradient descent and Markov chains
Aymeric Dieuleveut, Alain Durmus, and Francis Bach · 2020
Cited alongside, same era.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mezard, and Lenka Zdeborová · 2021
Later among the works it cites.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2022
Later among the works it cites.
Phase diagram of stochastic gradient descent in high-dimensional two-layer neural networks
Rodrigo Veiga, Ludovic STEPHAN, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborova · 2022
Later among the works it cites.
High-dimensional limit theorems for SGD: Effective dynamics and critical scaling
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2022
Later among the works it cites.
On the symmetries in the dynamics of wide two-layer neural networks, 2022
Karl Hajjar and Lenaic Chizat · 2022
Later among the works it cites.
Gradient flow dynamics of shallow reLU networks for square loss and orthogonal inputs
Etienne Boursier, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Later among the works it cites.
On Uniform Boundedness Properties of SGD and its Momentum Variants
Xiaoyu Wang and Mikael Johansson · 2022
Later among the works it cites.
Gradient descent on infinitely wide neural networks: Global convergence and generalization
Francis Bach and Lénaic Chizat · 2022
Later among the works it cites.