Fetching the paper…
Reading the bibliography…
Several recent works have shown separation results between deep neural networks, and hypothesis classes with inferior approximation capacity such as shallow networks or kernel classes.
Learning boolean circuits with neural networks
Eran Malach and Shai Shalev-Shwartz · 1910
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno, Vladimir Ya Lin, Allan Pinkus, and Shimon Schocken · 1993
Earlier work this paper cites.
Weakly learning dnf and characterizing statistical query learning using fourier analysis
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
A characterization of strong learnability in the statistical query model
Hans Ulrich Simon · 2007
Earlier work this paper cites.
Uniform approximation of functions with random bases
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Characterizing statistical query learning: simplified notions and proofs
Balázs Szörényi · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
A complete characterization of statistical query learning with applications to evolvability
Vitaly Feldman · 2012
Earlier work this paper cites.
Computational bounds on statistical query learning
Vitaly Feldman and Varun Kanade · 2012
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Cited alongside, same era.
Depth separation for neural networks
Amit Daniely · 2017
Cited alongside, same era.
Depth-width tradeoffs in approximating natural functions with neural networks
Itay Safran and Ohad Shamir · 2017
Cited alongside, same era.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Cited alongside, same era.
Distribution-specific hardness of learning neural networks
Ohad Shamir · 2018
Cited alongside, same era.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
Poly-time universality and limitations of deep learning
Emmanuel Abbe and Colin Sandon · 2020
Later among the works it cites.
Better depth-width trade-offs for neural networks through the lens of dynamical systems
Vaggos Chatziafratis, Sai Ganesh Nagarajan, and Ioannis Panageas · 2020
Later among the works it cites.
Learning parities with neural networks
Amit Daniely and Eran Malach · 2020
Later among the works it cites.
Agnostic learning of a single neuron with gradient descent
Spencer Frei, Yuan Cao, and Quanquan Gu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yitong Sun, Anna Gilbert, and Ambuj Tewari · 2018
Cited alongside, same era.
What can resnet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Cited alongside, same era.
Depth-width trade-offs for relu networks via sharkovsky’s theorem
Vaggos Chatziafratis, Sai Ganesh Nagarajan, Ioannis Panageas, and Xiao Wang · 2019
Cited alongside, same era.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Cited alongside, same era.
Depth separations in neural networks: What is actually being separated?
Itay Safran, Ronen Eldan, and Ohad Shamir · 2019
Cited alongside, same era.
Is deeper better only when shallow is good?
Eran Malach and Shai Shalev-Shwartz
Cited in the paper.
Superpolynomial lower bounds for learning one-layer neural networks using gradient descent
Surbhi Goel, Aravind Gollakota, Zhihan Jin, Sushrut Karmalkar, and Adam Klivans · 2020
Later among the works it cites.
Approximate is good enough: Probabilistic variants of dimensional and margin complexity
Pritish Kamath, Omar Montasser, and Nathan Srebro · 2020
Later among the works it cites.
Neural networks with small weights and depth-separation barriers
Gal Vardi and Ohad Shamir · 2020
Later among the works it cites.
Learning a single neuron with gradient methods
Gilad Yehudai and Ohad Shamir · 2020
Later among the works it cites.
Size and depth separation in approximating natural functions with neural networks
Gal Vardi, Daniel Reichman, Toniann Pitassi, and Ohad Shamir · 2021
Closest in time.