Fetching the paper…
Reading the bibliography…
We study the role of $L_2$ regularization in deep learning, and uncover simple relations between the performance of the model, the $L_2$ coefficient, the learning rate, and the number of training steps.
Learning distributed representations of concepts
G. E. Hinton · 1986
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Exploring Generalization in Deep Learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Earlier work this paper cites.
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
Levent Sagun, Utku Evci, V. Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
The implicit bias of gradient descent on separable data, 2017
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2017
Earlier work this paper cites.
L2 regularization versus batch and weight normalization, 2017
Twan van Laarhoven · 2017
Earlier work this paper cites.
Reconciling modern machine learning practice and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mand al · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne · 2018
Cited alongside, same era.
Gradient Descent Happens in a Tiny Subspace
Guy Gur-Ari, Daniel A. Roberts, and Ethan Dyer · 2018
Cited alongside, same era.
Norm matters: efficient and accurate normalization schemes in deep networks, 2018
Elad Hoffer, Ron Banner, Itay Golan, and Daniel Soudry · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A Hamprecht, Yoshua Bengio, and Aaron Courville · 2018
Cited alongside, same era.
An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Later among the works it cites.
An exponential learning rate schedule for deep learning, 2019
Zhiyuan Li and Sanjeev Arora · 2019
Later among the works it cites.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Yasaman Bahri, Jaehoon Lee, Greg Yang, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-dickstein · 2019
Later among the works it cites.
On the asymptotics of wide networks with polynomial activations, 2020
Kyle Aitken and Guy Gur-Ari · 2020
Closest in time.
Asymptotics of wide networks from feynman diagrams
Ethan Dyer and Guy Gur-Ari · 2020
Closest in time.
Scaling description of generalization with number of parameters in deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regularization matters: Generalization and optimization of neural nets v.s. their induced kernel, 2018
Colin Wei, Jason D. Lee, Qiang Liu, and Tengyu Ma · 2018
Cited alongside, same era.
Three mechanisms of weight decay regularization, 2018
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2018
Cited alongside, same era.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J. Zico Kolter, and Ryan J. Tibshirani · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism, 2020
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate
Zhiyuan Li, Kaifeng Lyu, and Sanjeev Arora · 2020
Closest in time.
Whitening and second order optimization both destroy information about the dataset, and can make generalization impossible, 2020
Neha S. Wadia, Daniel Duckworth, Samuel S. Schoenholz, Ethan Dyer, and Jascha Sohl-Dickstein · 2020
Closest in time.