Fetching the paper…
Reading the bibliography…
Understanding the implicit bias of training algorithms is of crucial importance in order to explain the success of overparametrised neural networks.
On strong solutions and explicit formulas for solutions of stochastic integral equations
Alexander Ju Veretennikov · 1981
Earlier work this paper cites.
Spectral gap and concentration for some spherically symmetric probability measures
Sergey G Bobkov · 2003
Earlier work this paper cites.
Fast global convergence of gradient methods for high-dimensional statistical recovery
Alekh Agarwal, Sahand Negahban, and Martin J. Wainwright · 2012
Earlier work this paper cites.
Brownian motion and stochastic calculus , volume 113
Ioannis Karatzas and Steven Shreve · 2012
Earlier work this paper cites.
Stability of stochastic differential equations
Rafail Khasminskii · 2012
Earlier work this paper cites.
Continuous martingales and Brownian motion , volume 293
Daniel Revuz and Marc Yor · 2013
Earlier work this paper cites.
Analysis and geometry of Markov diffusion operators , volume 103
Dominique Bakry, Ivan Gentil, Michel Ledoux, et al · 2014
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalisation
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Implicit regularisation in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Later among the works it cites.
Exponentiated gradient meets gradient descent
Udaya Ghai, Elad Hazan, and Yoram Singer · 2020
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrised models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
A continuous-time mirror descent approach to sparse phase retrieval
Fan Wu and Patrick Rebeschini · 2020
Later among the works it cites.
Label noise sgd provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D Lee · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations
Qianxiao Li, Cheng Tai, and Weinan E · 2019
Cited alongside, same era.
Implicit regularisation for optimal sparse recovery
Tomas Vaškevičius, Varun Kanade, and Patrick Rebeschini · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Martin J Wainwright · 2019
Cited alongside, same era.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Cited alongside, same era.
The implicit regularisation of stochastic gradient flow for least squares
Alnur Ali, Edgar Dobriban, and Ryan Tibshirani · 2020
Cited alongside, same era.
Implicit regularisation for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Cited alongside, same era.
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason Lee, and Tengyu Ma · 2021
Later among the works it cites.
Characterizing the implicit bias via a primal-dual analysis
Ziwei Ji and Matus Telgarsky · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Stochastic gradient descent with noise of machine learning type. Part I: Discrete time analysis, 2021
Stephan Wojtowytsch · 2021
Later among the works it cites.
What happens after sgd reaches zero loss? –a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2022
Closest in time.