Fetching the paper…
Reading the bibliography…
Overparameterization refers to the important phenomenon where the width of a neural network is chosen such that learning algorithms can provably attain zero loss in nonconvex training.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Incorporating second-order functional knowledge for better option pricing
Charles Dugas, Yoshua Bengio, François Bélisle, Claude Nadeau, and René Garcia · 2000
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S. Wright · 2006
Earlier work this paper cites.
NIST Handbook of Mathematical Functions Paperback and CD-ROM
Frank W. J. Olver, Daniel W. Lozier, Ronald F. Boisvert, and Charles W. Clark · 2010
Earlier work this paper cites.
Efficient BackProp. In Neural networks: Tricks of the Trade
Yann A. LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Introduction to the Non-asymptotic Analysis of Random Matrices
Roman Vershynin · 2012
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (ELUs)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Earlier work this paper cites.
From error bounds to the complexity of first-order descent methods for convex functions
Jérôme Bolte, Trong Phong Nguyen, Juan Peypouquet, and Bruce W Suter · 2017
Earlier work this paper cites.
Globally optimal gradient descent for a ConvNet with Gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Earlier work this paper cites.
Improved training of Wasserstein GANs
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Earlier work this paper cites.
Semi-supervised learning with GANs: Manifold invariance with improved inference
Abhishek Kumar, Prasanna Sattigeri, and Tom Fletcher · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
On the power of over-parametrization in neural networks with quadratic activation
Simon S. Du and Jason D. Lee · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Memorization precedes generation: Learning unsupervised GANs with memory networks
Youngjin Kim, Minjung Kim, and Gunhee Kim · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Nonconvex optimization meets low-rank matrix factorization: An overview
Yuejie Chi, Yue M Lu, and Yuxin Chen · 2019
Cited alongside, same era.
On lazy training in differentiable programming
On learning over-parameterized neural networks: A functional approximation perspective
Lili Su and Pengkun Yang · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Later among the works it cites.
Harnessing the power of infinitely wide deep nets on small-data tasks
Sanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang, and Dingli Yu · 2020
Later among the works it cites.
Neural networks learning and memorization with (almost) no over-parameterization
Amit Daniely · 2020
Later among the works it cites.
Training linear neural networks: Non-local convergence and complexity results
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
Limitations of lazy training of two-layers neural networks
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
Gradient descent finds global minima for generalizable deep neural networks of practical sizes
Kenji Kawaguchi and Jiaoyang Huang · 2019
Cited alongside, same era.
Fast and provable ADMM for learning with generative priors
Fabian Latorre, Armin Eftekhari, and Volkan Cevher · 2019
Cited alongside, same era.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
Armin Eftekhari · 2020
Later among the works it cites.
Gaussian error linear units (GELU)
Dan Hendrycks and Kevin Gimpel · 2020
Later among the works it cites.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks
Ziwei Ji and Matus Telgarsky · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak · 2020
Later among the works it cites.
A mean field analysis of deep resnet and beyond: Towards provably optimization via overparameterization from depth
Yiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu, and Lexing Ying · 2020
Later among the works it cites.
Global convergence of deep networks with one wide layer followed by pyramidal topology
Quynh Nguyen and Marco Mondelli · 2020
Later among the works it cites.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Empirical evaluation of rectified activations in convolutional network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2020
Later among the works it cites.
How much over-parameterization is sufficient to learn deep ReLU networks?
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2021
Closest in time.