Fetching the paper…
Reading the bibliography…
The success of neural networks over the past decade has established them as effective models for many relevant data generating processes.
On the uniform convergence of relative frequencies of events to their probabilities
V. N. Vapnik and A. Ya. Chervonenkis · 1971
Earlier work this paper cites.
Robust locally weighted regression and smoothing scatterplots
William S Cleveland · 1979
Earlier work this paper cites.
A theory of the learnable
Leslie G Valiant · 1984
Earlier work this paper cites.
Almost linear VC dimension bounds for piecewise polynomial networks
Peter Bartlett, Vitaly Maiorov, and Ron Meir · 1998
Earlier work this paper cites.
Algorithms and SQ lower bounds for PAC learning one-hidden-layer ReLU networks
Ilias Diakonikolas, Daniel M. Kane, Vasilis Kontonis, and Nikos Zarifis · 2006
Earlier work this paper cites.
Superpolynomial lower bounds for learning one-layer neural networks using gradient descent, 2020
Surbhi Goel, Aravind Gollakota, Zhihan Jin, Sushrut Karmalkar, and Adam Klivans · 2006
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Provable bounds for learning some deep representations
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Generalization bounds for neural networks through tensor factorization
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
The optimal sample complexity of PAC learning
Steve Hanneke · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
l1-regularized neural networks are improperly learnable in polynomial time
Yuchen Zhang, Jason D Lee, and Michael I Jordan · 2016
Cited alongside, same era.
Learning one-hidden-layer neural networks with landscape design
On the complexity of learning neural networks
Le Song, Santosh Vempala, John Wilmes, and Bo Xie · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Later among the works it cites.
Nearly-tight VC-dimension and pseudodimension bounds for piecewise linear neural networks
Peter L. Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
Leslie N. Smith · 2017
Cited alongside, same era.
Later among the works it cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Guaranteed recovery of one-hidden-layer neural networks via cross entropy
Haoyu Fu, Yuejie Chi, and Yingbin Liang · 2020
Later among the works it cites.
An information-theoretic framework for supervised learning, 2022
Hong Jun Jeon and Benjamin Van Roy · 2022
Closest in time.