Fetching the paper…
Reading the bibliography…
Gradient descent can be surprisingly good at optimizing deep neural networks without overfitting and without explicit regularization.
Error analysis of floating-point computation
James H. Wilkinson · 1960
Earlier work this paper cites.
Double backpropagation increasing generalization performance
Harris Drucker and Yann Le Cun · 1992
Earlier work this paper cites.
The life-span of backward error analysis for numerical integrators
Ernst Hairer and Christian Lubich · 1997
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Geometric numerical integration , volume 31
Ernst Hairer, Christian Lubich, and Gerhard Wanner · 2006
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Earlier work this paper cites.
Why does deep and cheap learning work so well?
Henry Lin and Max Tegmark · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Earlier work this paper cites.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D. Hoffman, and David M. Blei · 2017
Earlier work this paper cites.
The numerics of gans
Lars Mescheder, Sebastian Nowozin, and Andreas Geiger · 2017
Earlier work this paper cites.
Gradient descent gan optimization is locally stable
Vaishnavh Nagarajan and J Zico Kolter · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V. Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
Integration methods and optimization algorithms
Damien Scieur, Vincent Roulet, Francis Bach, and Alexandre d'Aspremont · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Leslie N. Smith · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Earlier work this paper cites.
The mechanics of n-player differentiable games
David Balduzzi, Sebastien Racaniere, James Martens, Jakob Foerster, Karl Tuyls, and Thore Graepel · 2018
Earlier work this paper cites.
Michael Betancourt, Michael I. Jordan, and Ashia C. Wilson · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne · 2018
Cited alongside, same era.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Later among the works it cites.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Mor Shpigel Nacson, Nathan Srebro, and Daniel Soudry · 2019
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Daniel S. Park, Jascha Sohl-dickstein, Quoc V. Le, and Sam Smith · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Just interpolate: Kernel "ridgeless" regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2018
Cited alongside, same era.
On the importance of single directions for generalization
Ari S. Morcos, David G.T. Barrett, Neil C. Rabinowitz, and Matthew Botvinick · 2018
Cited alongside, same era.
Generalization in deep networks: The role of distance from initialization
Vaishnavh Nagarajan and J. Zico Kolter · 2018
Cited alongside, same era.
Sgd implicitly regularizes generalization error
Daniel A. Roberts · 2018
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Sam Smith and Quoc V. Le · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Tomaso A. Poggio, Andrzej Banburski, and Qianli Liao · 2019
Later among the works it cites.
A type of generalization error induced by initialization in deep neural networks
Yaoyu Zhang, Zhiqin Xu, Tao Luo, and Zheng Ma · 2019
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2019
Later among the works it cites.
The implicit regularization of stochastic gradient flow for least squares
Alnur Ali, Edgar Dobriban, and Ryan Tibshirani · 2020
Closest in time.
Batch normalization biases residual blocks towards the identity function in deep networks
Soham De and Samuel Smith · 2020
Closest in time.
Uniform-in-time weak error analysis for stochastic gradient descent algorithms via diffusion approximation
Yuanyuan Feng, Tingran Gao, Lei Li, Jian-guo Liu, and Yulong Lu · 2020
Closest in time.
On dissipative symplectic integration with applications to gradient-based optimization
Guilherme França, Michael I. Jordan, and René Vidal · 2020
Closest in time.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Closest in time.
Haiku: Sonnet for JAX
Tom Hennigan, Trevor Cai, Tamara Norman, and Igor Babuschkin · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2020
Closest in time.
Chao Ma, Lei Wu, and Weinan E · 2020
Closest in time.
Training generative adversarial networks by solving ordinary differential equations
Chongli Qin, Yan Wu, Jost Tobias Springenberg, Andrew Brock, Jeff Donahue, Timothy P. Lillicrap, and Pushmeet Kohli · 2020
Closest in time.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Closest in time.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Closest in time.
The break-even point on optimization trajectories of deep neural networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2021
Closest in time.