Fetching the paper…
Reading the bibliography…
The success of deep networks has been attributed in part to their expressivity: per parameter, deep networks can approximate a richer class of functions than shallow networks.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
On the approximate realization of continuous mappings by neural networks
Ken-Ichi Funahashi · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1994
Earlier work this paper cites.
Almost linear VC dimension bounds for piecewise polynomial networks
Peter L Bartlett, Vitaly Maiorov, and Ron Meir · 1999
Earlier work this paper cites.
Approximation theory of the MLP model in neural networks
Allan Pinkus · 1999
Earlier work this paper cites.
An introduction to hyperplane arrangements
Richard P Stanley et al · 2004
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
On the complexity of neural network classifiers: A comparison between shallow and deep architectures
Monica Bianchini and Franco Scarselli · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Representation benefits of deep feedforward networks
Matus Telgarsky · 2015
Earlier work this paper cites.
On the expressive power of deep learning: A tensor analysis
Nadav Cohen, Or Sharir, and Amnon Shashua · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Cited alongside, same era.
Universal function approximation by deep neural nets with bounded width and ReLU activations
Boris Hanin · 2017
Cited alongside, same era.
Nearly-tight VC-dimension bounds for piecewise linear neural networks
Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2017
Cited alongside, same era.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Deep neural networks as Gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
The power of deeper networks for expressing natural functions
David Rolnick and Max Tegmark · 2018
Later among the works it cites.
Empirical bounds on linear regions of deep rectifier networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Henry W Lin, Max Tegmark, and David Rolnick · 2017
Cited alongside, same era.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Provable robustness of ReLU networks via maximization of linear regions
Francesco Croce, Maksym Andriushchenko, and Matthias Hein · 2018
Cited alongside, same era.
Deep neural networks are biased towards simple functions
Giacomo De Palma, Bobak Toussi Kiani, and Seth Lloyd · 2018
Cited alongside, same era.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Cited alongside, same era.
Products of many large random matrices and gradients in deep neural networks
Boris Hanin and Mihai Nica · 2018
Cited alongside, same era.
Thiago Serra and Srikumar Ramalingam · 2018
Later among the works it cites.
Bounding and counting linear regions of deep neural networks
Thiago Serra, Christian Tjandraatmadja, and Srikumar Ramalingam · 2018
Later among the works it cites.
Nearly-tight VC-dimension and pseudodimension bounds for piecewise linear neural networks
Peter L Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2019
Closest in time.
Asymptotics of wide networks from Feynman diagrams
Ethan Dyer and Guy Gur-Ari · 2019
Closest in time.
Complexity of linear regions in deep networks
Boris Hanin and David Rolnick · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.