Fetching the paper…
Reading the bibliography…
In this paper, we investigate the impact of stochasticity and large stepsizes on the implicit regularisation of gradient descent (GD) and stochastic gradient descent (SGD) over diagonal linear networks.
A stochastic approxiation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
L.M. Bregman · 1967
Earlier work this paper cites.
Stochastic Processes
J. L. Doob · 1990
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Stable signal recovery from incomplete and inaccurate measurements
E. Candès, J. Romberg, and T. Tao · 2006
Earlier work this paper cites.
A simple proof of the restricted isometry property for random matrices
R. Baraniuk, M. Davenport, R. DeVore, and M. Wakin · 2008
Earlier work this paper cites.
Concentration of measure
Terrence Tao · 2010
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Margins, shrinkage, and boosting
Matus Telgarsky · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
A generalized online mirror descent with applications to classification and regression
Francesco Orabona, Koby Crammer, and Nicolò Cesa-Bianchi · 2015
Earlier work this paper cites.
A variational analysis of stochastic gradient algorithms
Stephan Mandt, Matthew D. Hoffman, and David M. Blei · 2016
Earlier work this paper cites.
A descent lemma beyond Lipschitz gradient continuity: first-order methods revisited and applications
H. H Bauschke, J. Bolte, and M. Teboulle · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Train longer, generalize better: Closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Three factors influencing minima in SGD, 2017
Stanisław Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Width of minima reached by stochastic gradient descent is influenced by learning rate to batch size ratio
Stanislaw Jastrzkebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Earlier work this paper cites.
An alternative view: When does SGD escape local minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Cited alongside, same era.
Revisiting small batch training for deep neural networks
Dominic Masters and Carlo Luschi · 2018
Cited alongside, same era.
A Bayesian perspective on generalization and stochastic gradient descent
Samuel L. Smith and Quoc V. Le · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
High-Dimensional Probability: An Introduction with Applications in Data Science
Roman Vershynin · 2018
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Label noise SGD provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D. Lee · 2021
Later among the works it cites.
Fast stochastic Bregman gradient methods: Sharp analysis and variance reduction
Radu Alexandru Dragomir, Mathieu Even, and Hadrien Hendrikx · 2021
Later among the works it cites.
Concentration of non-isotropic random tensors with applications to learning and empirical risk minimization
Mathieu Even and Laurent Massoulie · 2021
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z. HaoChen, Colin Wei, Jason Lee, and Tengyu Ma · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
Rotem Mulayoff, Tomer Michaeli, and Daniel Soudry · 2021
Later among the works it cites.
Analysis of boolean functions, 2021
Ryan O’Donnell · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lei Wu, Chao Ma, and Weinan E · 2018
Cited alongside, same era.
On Lazy Training in Differentiable Programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
Fengxiang He, Tongliang Liu, and Dacheng Tao · 2019
Cited alongside, same era.
On the relation between the sharpest directions of DNN loss and the SGD step length
Stanisław Jastrzebski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amost Storkey · 2019
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Cited alongside, same era.
Later among the works it cites.
Implicit bias of SGD for diagonal linear networks: a provable benefit of stochasticity
S. Pesme, L. Pillaud-Vivien, and N. Flammarion · 2021
Later among the works it cites.
Stochastic gradient descent with noise of machine learning type. part II: Continuous time analysis
Stephan Wojtowytsch · 2021
Later among the works it cites.
Direction matters: On the implicit bias of stochastic gradient descent with moderate learning rate
Jingfeng Wu, Difan Zou, Vladimir Braverman, and Quanquan Gu · 2021
Later among the works it cites.
Learning threshold neurons via the "edge of stability"
Kwangjun Ahn, Sébastien Bubeck, Sinho Chewi, Yin Tat Lee, Felipe Suarez, and Yi Zhang · 2022
Later among the works it cites.
SGD with large step sizes learns sparse features
M. Andriushchenko, A. Varre, L. Pillaud-Vivien, and N. Flammarion · 2022
Later among the works it cites.
Incremental learning in diagonal linear networks
Raphaël Berthier · 2022
Later among the works it cites.
On the benefits of large learning rates for kernel methods
G. Beugnot, J. Mairal, and A. Rudi · 2022
Later among the works it cites.
On gradient descent convergence beyond the edge of stability, 2022
Lei Chen and Joan Bruna · 2022
Later among the works it cites.
Stochastic training is not necessary for generalization
Jonas Geiping, Micah Goldblum, Phillip E Pope, Michael Moeller, and Tom Goldstein · 2022
Later among the works it cites.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Later among the works it cites.
Label noise (stochastic) gradient descent implicitly solves the lasso for quadratic parametrisation
L. Pillaud-Vivien, J. Reygner, and N. Flammarion · 2022
Later among the works it cites.
Large learning rate tames homogeneity: Convergence and balancing effect
Yuqing Wang, Minshuo Chen, Tuo Zhao, and Molei Tao · 2022
Later among the works it cites.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D. Lee · 2023
Closest in time.
Saddle-to-saddle dynamics in diagonal linear networks
Scott Pesme and Nicolas Flammarion · 2023
Closest in time.
Understanding edge-of-stability training dynamics with a minimalist example
Xingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou, and Rong Ge · 2023
Closest in time.