Fetching the paper…
Reading the bibliography…
The generalization performance of a machine learning algorithm such as a neural network depends in a non-trivial way on the structure of the data distribution.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
On a Stochastic Approximation Method
K. L. Chung · 1954
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Asymptotic Properties of Statistical Estimators in Stochastic Programming
Alexander Shapiro · 1989
Earlier work this paper cites.
Learning processes in neural networks
Tom M. Heskes and Bert Kappen · 1991
Earlier work this paper cites.
Eigenvalues of covariance matrices: Application to neural-network learning
Yann LeCun, Ido Kanter, and Sara A. Solla · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. Polyak and A. Juditsky · 1992
Earlier work this paper cites.
On-line learning with a perceptron
M Biehl and P Riegler · 1994
Earlier work this paper cites.
Statistical mechanical analysis of the dynamics of learning in perceptrons
C. Mace and A. Coolen · 1998
Earlier work this paper cites.
Advanced Mathematical Methods for Scientists and Engineers: Asymptotic Methods and Perturbation Theory , volume 1
Carl Bender and Steven Orszag · 1999
Earlier work this paper cites.
Dynamics of on-line gradient descent learning for multilayer neural networks
David Saad and Sara Solla · 1999
Earlier work this paper cites.
Statistical Mechanics of Learning
A. Engel and C. Van den Broeck · 2001
Earlier work this paper cites.
Learning curves for stochastic gradient descent in linear feedforward networks
Justin Werfel, Xiaohui Xie, and H. Seung · 2003
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Carl Edward Rasmussen and Christopher K. I. Williams · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Online gradient descent learning algorithms
Yiming Ying and Massimiliano Pontil · 2008
Earlier work this paper cites.
From averaging to acceleration, there is only a step-size
Nicolas Flammarion and Francis Bach · 2015
Earlier work this paper cites.
Nonparametric stochastic approximation with large step-sizes
Aymeric Dieuleveut and Francis Bach · 2016
Earlier work this paper cites.
Harder, better, faster, stronger convergence rates for least-squares regression
Aymeric Dieuleveut, Nicolas Flammarion, and Francis Bach · 2016
Earlier work this paper cites.
Stochastic composite least-squares regression with convergence rate o ( 1 / n ) o(1/n)
Nicolas Flammarion and Francis Bach · 2017
Cited alongside, same era.
Deep learning scaling is predictable, empirically, 2017
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Cited alongside, same era.
Accelerating stochastic gradient descent for least squares regression
Prateek Jain, Sham M Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Cited alongside, same era.
Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes, 2018
Loucas Pillaud-Vivien, Alessandro Rudi, and Francis Bach · 2018
Cited alongside, same era.
Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification, 2020
Francesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborová · 2020
Later among the works it cites.
Neural tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2020
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Later among the works it cites.
Benign overfitting in ridge regression
Alexander Tsigler and Peter L Bartlett · 2020
Later among the works it cites.
An analysis of constant step size sgd in the non-convex regime: Asymptotic normality and bias, 2020
Lu Yu, Krishnakumar Balasubramanian, Stanislav Volgushev, and Murat A. Erdogdu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Cited alongside, same era.
Normal approximation for stochastic gradient descent via non-asymptotic rates of martingale clt
Andreas Anastasiou, Krishnakumar Balasubramanian, and Murat A. Erdogdu · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Cited alongside, same era.
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
Sebastian Goldt, Madhu Advani, Andrew M Saxe, Florent Krzakala, and Lenka Zdeborová · 2019
Cited alongside, same era.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George Dahl, Chris Shallue, and Roger B Grosse · 2019
Cited alongside, same era.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Cited alongside, same era.
Tight nonparametric convergence rates for stochastic gradient descent under the noiseless linear model, 2020
Raphaël Berthier, Francis Bach, and Pierre Gaillard · 2020
Cited alongside, same era.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Closest in time.
Risk bounds for over-parameterized maximum margin classification on sub-gaussian mixtures
Yuan Cao, Quanquan Gu, and Misha Belkin · 2021
Closest in time.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
Niladri S Chatterji and Philip M Long · 2021
Closest in time.
Asymptotic optimality in stochastic optimization
John C. Duchi and Feng Ruan · 2021
Closest in time.
The gaussian equivalence of generative models for learning with shallow neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2021
Closest in time.
The heavy-tail phenomenon in sgd
Mert Gurbuzbalaban, Umut Simsekli, and Lingjiong Zhu · 2021
Closest in time.
Capturing the learning curves of generic features maps for realistic data sets with a teacher-student model, 2021
Bruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2021
Closest in time.
Stochasticity helps to navigate rough landscapes: comparing gradient-descent-based algorithms in the phase retrieval problem, 2021
Francesca Mignacco, Pierfrancesco Urbani, and Lenka Zdeborová · 2021
Closest in time.
The deep bootstrap framework: Good online learners are good offline generalizers
Preetum Nakkiran, Behnam Neyshabur, and Hanie Sedghi · 2021
Closest in time.
The intrinsic dimension of images and its impact on learning
Phil Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum, and Tom Goldstein · 2021
Closest in time.
Last iterate convergence of sgd for least-squares in the interpolation regime, 2021
Aditya Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Closest in time.
Universal scaling laws in the gradient descent training of neural networks, 2021
Maksim Velikanov and Dmitry Yarotsky · 2021
Closest in time.
Benign overfitting of constant-stepsize sgd for linear regression
Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, and Sham M Kakade · 2021
Closest in time.