Fetching the paper…
Reading the bibliography…
An important characteristic of neural networks is their ability to learn representations of the input data with effective features for prediction, which is believed to be a key factor to their superior empirical performance.
Weakly learning dnf and characterizing statistical query learning using fourier analysis
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
B. Olshausen and D. Field · 1997
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Sparse coding and decorrelation in primary visual cortex during natural vision
William E Vinje and Jack L Gallant · 2000
Earlier work this paper cites.
Backward feature correction: How deep learning performs deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2001
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan · 2003
Earlier work this paper cites.
Feature purification: How adversarial training performs robust deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2005
Earlier work this paper cites.
Towards understanding hierarchical learning: Benefits of neural representations
Minshuo Chen, Yu Bai, Jason D Lee, Tuo Zhao, Huan Wang, Caiming Xiong, and Richard Socher · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2012
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Li Deng · 2012
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2012
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Ya Le and Xuan Yang · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
The comparative power of relu networks and polynomial kernels in the presence of sparse latent structure
Frederic Koehler and Andrej Risteski · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
What can resnet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Learning parities with neural networks
Amit Daniely and Eran Malach · 2020
Later among the works it cites.
Approximation schemes for relu regression
Ilias Diakonikolas, Surbhi Goel, Sushrut Karmalkar, Adam R Klivans, and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Training neural networks as learning data-adaptive kernels: Provable representation and approximation benefits
Xialiang Dou and Tengyuan Liang · 2020
Later among the works it cites.
Agnostic learning of a single neuron with gradient descent
Spencer Frei, Yuan Cao, and Quanquan Gu · 2020
Later among the works it cites.
Disentangling feature and lazy training in deep neural networks
Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart · 2020
Later among the works it cites.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Yu Bai and Jason D Lee · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Over parameterized two-level neural networks can learn near optimal feature representations, 2019
Cong Fang, Hanze Dong, and Tong Zhang · 2019
Cited alongside, same era.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Self-supervised visual feature learning with deep neural networks: A survey
Longlong Jing and Yingli Tian · 2020
Later among the works it cites.
Approximate is good enough: Probabilistic variants of dimensional and margin complexity
Pritish Kamath, Omar Montasser, and Nathan Srebro · 2020
Later among the works it cites.
Learning over-parametrized two-layer neural networks beyond ntk
Yuanzhi Li, Tengyu Ma, and Hongyang R Zhang · 2020
Later among the works it cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy · 2020
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
Learning a single neuron with gradient methods
Gilad Yehudai and Shamir Ohad · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Mathematical models of overparameterized neural networks
Cong Fang, Hanze Dong, and Tong Zhang · 2021
Later among the works it cites.
Quantifying the benefit of using differentiable learning over tangent kernels
Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro · 2021
Later among the works it cites.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed, 2021
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová · 2021
Later among the works it cites.
A local convergence theory for mildly over-parameterized two-layer neural network
Mo Zhou, Rong Ge, and Chi Jin · 2021
Later among the works it cites.