Fetching the paper…
Reading the bibliography…
While a lot of progress has been made in recent years, the dynamics of learning in deep nonlinear neural networks remain to this day largely misunderstood.
Ockham’s Razor: A Historical and Philosophical Analysis of Ockham’s Principle of Parsimony
Roger Ariew · 1976
Earlier work this paper cites.
Ordinary Differential Equations : An Elementary Textbook for Students of Mathematics, Engineering, and the Sciences
Morris Tenenbaum and Harry Pollard · 1985
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
On-line learning processes in artificial neural networks, 1993
Tom M. Heskes and Bert Kappen · 1993
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Statistical Learning Theory
Vladimir Naumovich Vapnik · 1998
Earlier work this paper cites.
General conditions for predictivity in learning theory
Tomaso Poggio, Ryan Rifkin, Sayan Mukherjee, and Partha Niyogi · 2004
Earlier work this paper cites.
Are loss functions all the same?
Lorenzo Rosasco, Ernesto De Vito, Andrea Caponnetto, Michele Piana, and Alessandro Verri · 2004
Earlier work this paper cites.
The loss surface of multilayer networks
Anna Choromanska, Mikael Henaff, Michaël Mathieu, Gérard Ben Arous, and Yann LeCun · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization, 2014
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey E. Hinton · 2015
Cited alongside, same era.
Deep Linear Neural Networks: A Theory of Learning in the Brain and Mind
Andrew Michael Saxe · 2015
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Closest in time.
Dogs vs. cats, 2018
Kaggle · 2018
Closest in time.
An alternative view: When does sgd escape local minima?
Robert Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Closest in time.
The dynamics of learning: A random matrix approach
Zhenyu Liao and Romain Couillet · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S. Advani and Andrew M. Saxe · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel S. Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
Svcca: Singular vector canonical correlation analysis for deep understanding and improvement
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, and Nathan Srebro · 2017
Cited alongside, same era.
Learning hierarchical categories in deep neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli
Cited in the paper.
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Closest in time.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Nathan Srebro, and Daniel Soudry · 2018
Closest in time.
Is generator conditioning causally related to gan performance?
Augustus Odena, Jacob Buckman, Catherine Olsson, Tom B. Brown, Christopher Olah, Colin Raffel, and Ian J. Goodfellow · 2018
Closest in time.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle Perez, Chico Q. Camargo, and Ard A. Louis · 2018
Closest in time.
Exponential integral, 2018
Wiki · 2018
Closest in time.
Convergence of sgd in learning relu models with separable data
Tengyu Xu, Yi Zhou, Kaiyi Ji, and Yingbin Liang · 2018
Closest in time.