Fetching the paper…
Reading the bibliography…
We empirically evaluate common assumptions about neural networks that are widely held by practitioners and theorists alike.
On the effect of low-rank weights on adversarial robustness of neural networks
Peter Langenberg, Emilio Rafael Balda, Arash Behboodi, and Rudolf Mathar · 1901
Earlier work this paper cites.
Fixup Initialization: Residual Learning Without Normalization
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 1901
Earlier work this paper cites.
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 1902
Earlier work this paper cites.
Methods of conjugate gradients for solving linear systems , volume 49
Magnus Rudolph Hestenes and Eduard Stiefel · 1952
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o (1/kˆ 2)
Yurii Nesterov · 1983
Earlier work this paper cites.
The strength of weak learnability
Robert E Schapire · 1990
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Kernel principal component analysis
Bernhard Schölkopf, Alexander Smola, and Klaus-Robert Müller · 1997
Earlier work this paper cites.
Least squares support vector machine classifiers
Johan AK Suykens and Joos Vandewalle · 1999
Earlier work this paper cites.
The Loss Surfaces of Multilayer Networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J. Goodfellow, Oriol Vinyals, and Andrew M. Saxe · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Xception: Deep Learning with Depthwise Separable Convolutions
François Chollet · 2016
Earlier work this paper cites.
Deep Learning without Poor Local Minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Local minima in training of neural networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
The Shattered Gradients Problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, J. P. Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Cited alongside, same era.
Sharp Minima Can Generalize For Deep Nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Global Optimality in Neural Network Training
B. D. Haeffele and R. Vidal · 2017
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Later among the works it cites.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James Brecht · 2018
Later among the works it cites.
Understanding the Loss Surface of Neural Networks for Binary Classification
Shiyu Liang, Ruoyu Sun, Yixuan Li, and R. Srikant · 2018
Later among the works it cites.
On the loss landscape of a class of deep neural networks with no bad local valleys
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2017
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Deep Neural Networks as Gaussian Processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Cited alongside, same era.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Cited alongside, same era.
Spurious Local Minima are Common in Two-Layer ReLU Neural Networks
Itay Safran and Ohad Shamir · 2017
Cited alongside, same era.
L2 Regularization versus Batch and Weight Normalization
Twan van Laarhoven · 2017
Cited alongside, same era.
Quynh Nguyen, Mahesh Chandra Mukkamala, and Matthias Hein · 2018
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Later among the works it cites.
How Does Batch Normalization Help Optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Later among the works it cites.
The singular values of convolutional layers
Hanie Sedghi, Vineet Gupta, and Philip M Long · 2018
Later among the works it cites.
Minimum norm solutions do not always generalize well for over-parameterized problems
Vatsal Shah, Anastasios Kyrillidis, and Sujay Sanghavi · 2018
Later among the works it cites.
Small nonlinearities in activation functions create bad local minima in neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Later among the works it cites.
Three Mechanisms of Weight Decay Regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2018
Later among the works it cites.
Gradient Descent Finds Global Minima of Deep Neural Networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Closest in time.
Understanding generalization through visualizations
W Ronny Huang, Zeyad Emam, Micah Goldblum, Liam Fowl, Justin K Terry, Furong Huang, and Tom Goldstein · 2019
Closest in time.
Small nonlinearities in activation functions create bad local minima in neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2019
Closest in time.
Piecewise linear activations substantially shape the loss surfaces of neural networks
Fengxiang He, Bohan Wang, and Dacheng Tao · 2020
Closest in time.