Fetching the paper…
Reading the bibliography…
We prove that the binary classifiers of bit strings generated by random wide deep neural networks with ReLU activation function are biased towards simple functions.
On the complexity of finite sequences
Abraham Lempel and Jacob Ziv · 1976
Earlier work this paper cites.
A universal algorithm for sequential data compression
J. Ziv and A. Lempel · 1977
Earlier work this paper cites.
Modeling by shortest data description
Jorma Rissanen · 1978
Earlier work this paper cites.
Occam’s razor
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K Warmuth · 1987
Earlier work this paper cites.
What size net gives valid generalization?
Eric B Baum and David Haussler · 1989
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
Discovering neural nets with low kolmogorov complexity and high generalization capability
Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Computing with infinite networks
Christopher KI Williams · 1997
Earlier work this paper cites.
On tables of random numbers
Andrei N Kolmogorov · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
The mnist database of handwritten digits, http://yann.lecun.com/exdb/mnist/, 1998
Yann LeCun, Corinna Cortes, and Christopher J.C. Burges · 1998
Earlier work this paper cites.
Some pac-bayesian theorems
David A McAllester · 1999
Earlier work this paper cites.
Extreme values of random processes
Anton Bovier · 2005
Earlier work this paper cites.
Generalization ability of boolean functions implemented in feedforward neural networks
Leonardo Franco · 2006
Earlier work this paper cites.
Pac-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning
O. Catoni · 2007
Earlier work this paper cites.
Multidimensional Diffusion Processes
Daniel W. Stroock and S. R. Srinivasa Varadhan · 2007
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 2012
Earlier work this paper cites.
The Nature of Statistical Learning Theory
V.N. Vapnik · 2013
Earlier work this paper cites.
Tighter pac-bayes bounds through distribution-dependent priors
Guy Lever, François Laviolette, and John Shawe-Taylor · 2013
Earlier work this paper cites.
An Introduction to Kolmogorov Complexity and Its Applications
M. Li and P. Vitanyi · 2013
Earlier work this paper cites.
No free lunch versus occam’s razor in supervised learning
Tor Lattimore and Marcus Hutter · 2013
Earlier work this paper cites.
On the number of response regions of deep feed forward networks with piece-wise linear activations
Razvan Pascanu, Guido Montufar, and Yoshua Bengio · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Cited alongside, same era.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, and Nathan Srebro · 2017
Later among the works it cites.
Lower bounds on the robustness to adversarial perturbations
Jonathan Peck, Joris Roels, Bart Goossens, and Yvan Saeys · 2017
Later among the works it cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Later among the works it cites.
Sensitivity and generalization in neural networks: an empirical study
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Open problem: The landscape of the loss surfaces of multilayer networks
Anna Choromanska, Yann LeCun, and Gérard Ben Arous · 2015
Cited alongside, same era.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Cited alongside, same era.
An analysis of deep neural network models for practical applications
Alfredo Canziani, Adam Paszke, and Eugenio Culurciello · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Closest in time.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Closest in time.
Data-dependent pac-bayes priors via differential privacy
Gintare Karolina Dziugaite and Daniel M Roy · 2018
Closest in time.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Closest in time.
On the importance of single directions for generalization
Ari S Morcos, David GT Barrett, Neil C Rabinowitz, and Matthew Botvinick · 2018
Closest in time.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Closest in time.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Laurence Aitchison, and Carl Edward Rasmussen · 2018
Closest in time.
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel S Schoenholz, and Jeffrey Pennington · 2018
Closest in time.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Closest in time.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2018
Closest in time.
The relationship between pac, the statistical physics framework, the bayesian framework, and the vc framework
David H Wolpert · 2018
Closest in time.
Input–output maps are strongly biased towards simple outputs
Kamaludin Dingle, Chico Q Camargo, and Ard A Louis · 2018
Closest in time.
Peter Hinz and Sara van de Geer · 2018
Closest in time.
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2018
Closest in time.
Comparing dynamics: Deep neural networks versus glassy systems
Marco Baity-Jesi, Levent Sagun, Mario Geiger, Stefano Spigler, G Ben Arous, Chiara Cammarota, Yann LeCun, Matthieu Wyart, and Giulio Biroli · 2018
Closest in time.
Loss surface of xor artificial neural networks
Dhagash Mehta, Xiaojun Zhao, Edgar A Bernal, and David J Wales · 2018
Closest in time.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle-Perez, Chico Q. Camargo, and Ard A. Louis · 2019
Closest in time.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Yasaman Bahri, Jaehoon Lee, Greg Yang, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-dickstein · 2019
Closest in time.
Greg Yang · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
Understanding priors in bayesian neural networks at the unit level
Mariia Vladimirova, Jakob Verbeek, Pablo Mesejo, and Julyan Arbel · 2019
Closest in time.
Adversarial robustness may be at odds with simplicity
Preetum Nakkiran · 2019
Closest in time.