Fetching the paper…
Reading the bibliography…
This paper identifies a structural property of data distributions that enables deep neural networks to learn hierarchically.
Weakly learning dnf and characterizing statistical query learning using fourier analysis
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Bayesian network classifiers
Nir Friedman, Dan Geiger, and Moises Goldszmidt · 1997
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Clustering methods
Lior Rokach and Oded Maimon · 2005
Earlier work this paper cites.
Learning and smoothed analysis
Adam Tauman Kalai, Alex Samorodnitsky, and Shang-Hua Teng · 2009
Earlier work this paper cites.
An introduction to hierarchical linear modeling
Heather Woltman, Andrea Feldstain, J Christine MacKay, and Meredith Rocchi · 2012
Earlier work this paper cites.
Invariant scattering convolution networks
Joan Bruna and Stéphane Mallat · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
A probabilistic theory of deep learning
Ankit B Patel, Tan Nguyen, and Richard G Baraniuk · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Song Han, Jeff Pool, John Tran, and William J Dally · 2015
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Earlier work this paper cites.
Deep learning and hierarchal generative models
Elchanan Mossel · 2016
Earlier work this paper cites.
Learning functions: when is deep better than shallow
Hrushikesh Mhaskar, Qianli Liao, and Tomaso Poggio · 2016
Cited alongside, same era.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
Greedy layerwise learning can scale to imagenet
Eugene Belilovsky, Michael Eickenberg, and Edouard Oyallon · 2019
Later among the works it cites.
Training neural networks with local error signals
Arild Nøkland and Lars Hiller Eidnes · 2019
Later among the works it cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Later among the works it cites.
On the universality of deep learning
Emmanuel Abbe and Colin Sandon · 2020
Later among the works it cites.
The implications of local correlation on learning some deep functions
Eran Malach and Shai Shalev-Shwartz · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep convolutional framelets: A general deep learning framework for inverse problems
Jong Chul Ye, Yoseob Han, and Eunju Cha · 2018
Cited alongside, same era.
A provably correct algorithm for deep learning that actually works
Eran Malach and Shai Shalev-Shwartz · 2018
Cited alongside, same era.
The convergence rate of neural networks for learned functions of different frequencies
Ronen Basri, David Jacobs, Yoni Kasten, and Shira Kritchman · 2019
Cited alongside, same era.
Is deeper better only when shallow is good?
Eran Malach and Shai Shalev-Shwartz · 2019
Cited alongside, same era.
What can resnet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Cited alongside, same era.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Cited alongside, same era.
Later among the works it cites.
Towards understanding hierarchical learning: Benefits of neural representations
Minshuo Chen, Yu Bai, Jason D Lee, Tuo Zhao, Huan Wang, Caiming Xiong, and Richard Socher · 2020
Later among the works it cites.
Backward feature correction: How deep learning performs deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2020
Later among the works it cites.
Deep equals shallow for relu networks in kernel regimes
Alberto Bietti and Francis Bach · 2020
Later among the works it cites.
Why do deep residual networks generalize better than deep feedforward networks?—a neural tangent kernel perspective
Kaixuan Huang, Yuqing Wang, Molei Tao, and Tuo Zhao · 2020
Later among the works it cites.
Id3 learns juntas for smoothed product distributions
Alon Brutzkus, Amit Daniely, and Eran Malach · 2020
Later among the works it cites.
Decoupled greedy learning of cnns
Eugene Belilovsky, Michael Eickenberg, and Edouard Oyallon · 2020
Later among the works it cites.
On the power of differentiable learning versus pac and sq learning
Emmanuel Abbe, Pritish Kamath, Eran Malach, Colin Sandon, and Nathan Srebro · 2021
Closest in time.
Quantifying the benefit of using differentiable learning over tangent kernels
Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro · 2021
Closest in time.