Fetching the paper…
Reading the bibliography…
This work studies the behavior of shallow ReLU networks trained with the logistic loss via gradient descent on binary classification data where the underlying data distribution is general, and the (optimal) Bayes risk is not necessarily zero.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 1906
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
On the approximate realization of continuous mappings by neural networks
K. Funahashi · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R. Barron · 1993
Earlier work this paper cites.
Strong universal consistency of neural network classifiers
A. Farago and G. Lugosi · 1993
Earlier work this paper cites.
Approximation and learning of convex superpositions
Leonid Gurvits and Pascal Koiran · 1995
Earlier work this paper cites.
A probabilistic theory of pattern recognition
L. Devroye, L. Györfi, and G. Lugosi · 1996
Earlier work this paper cites.
Real analysis: modern techniques and their applications
Gerald B. Folland · 1999
Earlier work this paper cites.
Local operator theory, random matrices and Banach spaces
Kenneth R Davidson and Stanislaw J Szarek · 2001
Earlier work this paper cites.
Foundations of modern probability
Olav Kallenberg · 2002
Earlier work this paper cites.
Statistical behavior and consistency of classification methods based on convex risk minimization
Tong Zhang · 2004
Earlier work this paper cites.
Boosting with early stopping: Convergence and consistency
Tong Zhang and Bin Yu · 2005
Earlier work this paper cites.
Convexity, classification, and risk bounds
Peter L. Bartlett, Michael I. Jordan, and Jon D. McAuliffe · 2006
Earlier work this paper cites.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2006
Earlier work this paper cites.
AdaBoost is consistent
Peter L. Bartlett and Mikhail Traskin · 2007
Earlier work this paper cites.
Distributional generalization: A new kind of generalization
Preetum Nakkiran and Yamini Bansal · 2009
Earlier work this paper cites.
Tight hardness results for training depth-2 relu networks
Surbhi Goel, Adam R. Klivans, Pasin Manurangsi, and Daniel Reichman · 2011
Earlier work this paper cites.
Boosting: Foundations and Algorithms
Robert E. Schapire and Yoav Freund · 2012
Cited alongside, same era.
Boosting with the logistic loss is consistent
Matus Telgarsky · 2013
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Cited alongside, same era.
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Later among the works it cites.
A Note on Lazy Training in Supervised Differentiable Programming
Lénaïc Chizat and Francis Bach · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
Atsushi Nitanda and Taiji Suzuki · 2019
Later among the works it cites.
Samet Oymak and Mahdi Soltanolkotabi · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M. Roy · 2017
Cited alongside, same era.
On calibration of modern neural networks, 2017
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
Nonparametric regression using deep neural networks with relu activation function
Johannes Schmidt-Hieber · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Cited alongside, same era.
Later among the works it cites.
Foundations of Data Science
Avrim Blum, John Hopcroft, and Ravindran Kannan · 2020
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Later among the works it cites.
Agnostic learning of a single neuron with gradient descent
Spencer Frei, Yuan Cao, and Quanquan Gu · 2020
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Later among the works it cites.
Gradient methods never overfit on separable data
Ohad Shamir · 2020
Later among the works it cites.
Learning a single neuron with gradient methods
Gilad Yehudai and Ohad Shamir · 2020
Later among the works it cites.
Yu Bai, Song Mei, Huan Wang, and Caiming Xiong · 2021
Closest in time.
The smoking gun: Statistical theory improves neural network estimates
Alina Braun, Michael Kohler, Sophie Langer, and Harro Walk · 2021
Closest in time.
How much over-parameterization is sufficient to learn deep relu networks?
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2021
Closest in time.
Achieving small test error in mildly overparameterized neural networks
Shiyu Liang, Ruoyu Sun, and R. Srikant · 2021
Closest in time.
Parameter-free stochastic optimization of variationally coherent functions
Francesco Orabona and Dávid Pál · 2021
Closest in time.
Dominic Richards and Ilja Kuzborskij · 2021
Closest in time.