Fetching the paper…
Reading the bibliography…
Depth separation results propose a possible theoretical explanation for the benefits of deep neural networks over shallower architectures, establishing that the former possess superior approximation capabilities.
Tables for computing bivariate normal probabilities
Donald B Owen · 1956
Earlier work this paper cites.
A table of normal integrals: A table
Donald Bruce Owen · 1980
Earlier work this paper cites.
Exponentially many local minima for single neurons
Peter Auer, Mark Herbster, Manfred K Warmuth, et al · 1996
Earlier work this paper cites.
Foundations of Cryptography, Volume 2
Oded Goldreich · 2003
Earlier work this paper cites.
Theory of classification: A survey of some recent advances
Stéphane Boucheron, Olivier Bousquet, and Gábor Lugosi · 2005
Earlier work this paper cites.
The geometry of logconcave functions and sampling algorithms
László Lovász and Santosh Vempala · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Log-concavity and strong log-concavity: a review
Adrien Saumard and Jon A Wellner · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
On the quality of the initial basin in overspecified neural networks
Itay Safran and Ohad Shamir · 2016
Earlier work this paper cites.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Cited alongside, same era.
Depth separation for neural networks
Amit Daniely · 2017
Cited alongside, same era.
Why deep neural networks for function approximation?
Shiyu Liang and R Srikant · 2017
Cited alongside, same era.
Why and when can deep-but not shallow-networks avoid the curse of dimensionality: a review
Tomaso Poggio, Hrushikesh Mhaskar, Lorenzo Rosasco, Brando Miranda, and Qianli Liao · 2017
Cited alongside, same era.
Depth-width tradeoffs in approximating natural functions with neural networks
Itay Safran and Ohad Shamir · 2017
Cited alongside, same era.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2017
Cited alongside, same era.
Depth separations in neural networks: what is actually being separated?
Itay Safran, Ronen Eldan, and Ohad Shamir · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
Backward feature correction: How deep learning performs deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2020
Later among the works it cites.
Towards understanding hierarchical learning: Benefits of neural representations
Minshuo Chen, Yu Bai, Jason D Lee, Tuo Zhao, Huan Wang, Caiming Xiong, and Richard Socher · 2020
Later among the works it cites.
Agnostic learning of a single neuron with gradient descent
Spencer Frei, Yuan Cao, and Quanquan Gu · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lectures on convex optimization , volume 137
Yurii Nesterov et al · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Cited alongside, same era.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Yu Bai and Jason D Lee · 2019
Cited alongside, same era.
Learning neural networks with two nonlinear layers in polynomial time
Surbhi Goel and Adam R Klivans · 2019
Cited alongside, same era.
Neural tangent kernels, transportation mappings, and universal approximation
Ziwei Ji, Matus Telgarsky, and Ruicheng Xian · 2019
Cited alongside, same era.
Is deeper better only when shallow is good?
Eran Malach and Shai Shalev-Shwartz · 2019
Cited alongside, same era.
Later among the works it cites.
Neural networks with small weights and depth-separation barriers
Gal Vardi and Ohad Shamir · 2020
Later among the works it cites.
Learning a single neuron with gradient methods
Gilad Yehudai and Ohad Shamir · 2020
Later among the works it cites.
NIST Digital Library of Mathematical Functions
DLMF · 2021
Closest in time.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Closest in time.
On the approximation power of two-layer networks of random relus
Daniel Hsu, Clayton Sanford, Rocco A Servedio, and Emmanouil-Vasileios Vlatakis-Gkaragkounis · 2021
Closest in time.
The connection between approximation, depth separation and learnability in neural networks
Eran Malach, Gilad Yehudai, Shai Shalev-Shwartz, and Ohad Shamir · 2021
Closest in time.
Depth separation beyond radial functions
Luca Venturi, Samy Jelassi, Tristan Ozuch, and Joan Bruna · 2021
Closest in time.