Fetching the paper…
Reading the bibliography…
The early phase of training a deep neural network has a dramatic effect on the local curvature of the loss function.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
Michael F Hutchinson · 1990
Earlier work this paper cites.
Improving generalization performance using double backpropagation
H. Drucker and Y. Le Cun · 1992
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Contractive auto-encoders: Explicit invariance during feature extraction
Salah Rifai, Pascal Vincent, Xavier Muller, Xavier Glorot, and Yoshua Bengio · 2011
Earlier work this paper cites.
Efficient BackProp
Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Deep learning
Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Y. Le and X. Yang · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron C. Courville, Yoshua Bengio, and Simon Lacoste-Julien · 2017
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Levent Sagun, and Riccardo Zecchina · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martín Arjovsky, Vincent Dumoulin, and Aaron C. Courville · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Three Factors Influencing Minima in SGD
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos J. Storkey · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Earlier work this paper cites.
Implicit regularization in deep learning
Behnam Neyshabur · 2017
Earlier work this paper cites.
Spectral norm regularization for improving the generalizability of deep learning, 2017
Yuichi Yoshida and Takeru Miyato · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Understanding batch normalization
Johan Bjorck, Carla P. Gomes, Bart Selman, and Kilian Q. Weinberger · 2018
Earlier work this paper cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace, 2018
Guy Gur-Ari, Daniel A. Roberts, and Ethan Dyer · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry P. Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Clément Hongler, and Franck Gabriel · 2018
Cited alongside, same era.
Theory of deep learning iii: explaining the non-overfitting puzzle, 2018
Tomaso Poggio, Kenji Kawaguchi, Qianli Liao, Brando Miranda, Lorenzo Rosasco, Xavier Boix, Jack Hidary, and Hrushikesh Mhaskar · 2018
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Residual learning without normalization via better initialization
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 2019
Later among the works it cites.
Gradient ℓ 1 \ell_{1} regularization for quantization robustness
Milad Alizadeh, Arash Behboodi, Mart van Baalen, Christos Louizos, Tijmen Blankevoort, and Max Welling · 2020
Closest in time.
Jacobian adversarially regularized networks for robustness
Alvin Chan, Yi Tay, Yew-Soon Ong, and Jie Fu · 2020
Closest in time.
Coherent gradients: An approach to understanding generalization in gradient descent-based optimization
Satrajit Chatterjee · 2020
Closest in time.
Batch normalization biases residual blocks towards the identity function in deep networks
Soham De and Samuel L. Smith · 2020
Closest in time.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Samuel L. Smith and Quoc V. Le · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, and Nathan Srebro · 2018
Cited alongside, same era.
Gradient regularization improves accuracy of discriminative models, 2018
Dániel Varga, Adrián Csiszárik, and Zsolt Zombori · 2018
Cited alongside, same era.
Smoothout: Smoothing out sharp minima to improve generalization in deep learning, 2018
Wei Wen, Yandan Wang, Feng Yan, Cong Xu, Chunpeng Wu, Yiran Chen, and Hai Li · 2018
Cited alongside, same era.
Understanding training and generalization in deep learning by fourier analysis, 2018
Zhiqin John Xu · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
Critical learning periods in deep networks
Alessandro Achille, Matteo Rovere, and Stefano Soatto · 2019
Cited alongside, same era.
Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M. Roy, and Surya Ganguli · 2020
Closest in time.
The early phase of neural network training
Jonathan Frankle, David J. Schwab, and Ari S. Morcos · 2020
Closest in time.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann · 2020
Closest in time.
The surprising simplicity of the early-time learning dynamics of neural networks
Wei Hu, Lechao Xiao, Ben Adlam, and Jeffrey Pennington · 2020
Closest in time.
The break-even point on optimization trajectories of deep neural networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Closest in time.
Beyond synthetic noise: Deep learning on controlled noisy labels
Lu Jiang, Di Huang, Mason Liu, and Weilong Yang · 2020
Closest in time.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2020
Closest in time.
The two regimes of deep network training, 2020
Guillaume Leclerc and Aleksander Madry · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism, 2020
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
Dividemix: Learning with noisy labels as semi-supervised learning
Junnan Li, Richard Socher, and Steven C. H. Hoi · 2020
Closest in time.
Early-learning regularization prevents memorization of noisy labels
Sheng Liu, Jonathan Niles-Weed, Narges Razavian, and Carlos Fernandez-Granda · 2020
Closest in time.
New insights and perspectives on the natural gradient method, 2020
James Martens · 2020
Closest in time.
Learning from failure: Training debiased classifier from biased classifier
Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin · 2020
Closest in time.
Robust and on-the-fly dataset denoising for image classification
Jiaming Song, Yann Dauphin, Michael Auli, and Tengyu Ma · 2020
Closest in time.
On the interplay between noise and curvature and its effect on optimization and generalization
Valentin Thomas, Fabian Pedregosa, Bart van Merriënboer, Pierre-Antoine Manzagol, Yoshua Bengio, and Nicolas Le Roux · 2020
Closest in time.
Normalized flat minima: Exploring scale invariant definition of flat minima for neural networks using pac-bayesian analysis
Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama · 2020
Closest in time.
Implicit gradient regularization
David Barrett and Benoit Dherin · 2021
Closest in time.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Closest in time.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Closest in time.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David Barrett, and Soham De · 2021
Closest in time.