Fetching the paper…
Reading the bibliography…
Understanding the implicit bias of training algorithms is of crucial importance in order to explain the success of overparametrised neural networks.
Stochastic differential equations
Peter E Kloeden and Eckhard Platen · 1992
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
Bounds on the lambert function and their application to the outage analysis of user cooperation
Ioannis Chatzigeorgiou · 2013
Earlier work this paper cites.
Continuous martingales and Brownian motion , volume 293
Daniel Revuz and Marc Yor · 2013
Earlier work this paper cites.
A variational analysis of stochastic gradient algorithms
Stephan Mandt, Matthew D. Hoffman, and David M. Blei · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
A descent lemma beyond lipschitz gradient continuity: first-order methods revisited and applications
Heinz H Bauschke, Jérôme Bolte, and Marc Teboulle · 2017
Earlier work this paper cites.
Train longer, generalize better: Closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Three factors influencing minima in SGD
Stanislaw Jastrzebski, Zac Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Amos Storkey, and Yoshua Bengio · 2018
Cited alongside, same era.
An alternative view: When does SGD escape local minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Later among the works it cites.
Stochastic gradient and Langevin processes
Xiang Cheng, Dong Yin, Peter Bartlett, and Michael Jordan · 2020
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Later among the works it cites.
Exponentiated gradient meets gradient descent
Udaya Ghai, Elad Hazan, and Yoram Singer · 2020
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason D Lee, and Tengyu Ma · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
Fengxiang He, Tongliang Liu, and Dacheng Tao · 2019
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations
Qianxiao Li, Cheng Tai, and Weinan E · 2019
Cited alongside, same era.
Implicit regularization for optimal sparse recovery
Tomas Vaškevičius, Varun Kanade, and Patrick Rebeschini · 2019
Cited alongside, same era.
Steven R. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon · 2020
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Later among the works it cites.
The statistical complexity of early-stopped mirror descent
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
A continuous-time mirror descent approach to sparse phase retrieval
Fan Wu and Patrick Rebeschini · 2020
Later among the works it cites.
On the implicit bias of initialization shape: Beyond infinitesimal mirror descent
Shahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake Woodworth, Nathan Srebro, Amir Globerson, and Daniel Soudry · 2021
Closest in time.
Last iterate convergence of sgd for least-squares in the interpolation regime
Aditya Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Closest in time.
Stochastic gradient descent with noise of machine learning type. Part I: Discrete time analysis, 2021
Stephan Wojtowytsch · 2021
Closest in time.