Fetching the paper…
Reading the bibliography…
We establish conditions under which gradient descent applied to fixed-width deep networks drives the logistic loss to zero, and prove bounds on the rate of convergence.
“Learning polynomials with neural networks”
Alexandr Andoni, Rina Panigrahy, Gregory Valiant and Li Zhang · 1916
Earlier work this paper cites.
“Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition”
Thomas Cover · 1965
Earlier work this paper cites.
“Introduction to algorithms”
Thomas Cormen, Charles Leiserson, Ronald Rivest and Clifford Stein · 2009
Earlier work this paper cites.
“Convex optimization: algorithms and complexity”
Sébastien Bubeck · 2015
Earlier work this paper cites.
“In search of the real inductive bias: On the role of implicit regularization in deep learning”
Behnam Neyshabur, Ryota Tomioka and Nathan Srebro · 2015
Earlier work this paper cites.
“Convergence analysis of two-layer neural networks with ReLU activation”
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht and Oriol Vinyals · 2017
Earlier work this paper cites.
“On the learnability of fully-connected neural networks”
Yuchen Zhang, Jason Lee, Martin Wainwright and Michael Jordan · 2017
Earlier work this paper cites.
“Recovery Guarantees for One-hidden-layer Neural Networks”
Kai Zhong, Zhao Song, Prateek Jain, Peter Bartlett and Inderjit. Dhillon · 2017
Earlier work this paper cites.
“Overfitting or perfect fitting? Risk bounds for classification and regression rules that interpolate”
Mikhail Belkin, Daniel Hsu and Partha Mitra · 2018
Earlier work this paper cites.
“SGD learns over-parameterized networks that provably generalize on linearly separable data”
Alon Brutzkus, Amir Globerson, Eran Malach and Shai Shalev-Shwartz · 2018
Earlier work this paper cites.
“On the global convergence of gradient descent for over-parameterized models using optimal transport”
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
“Gradient descent provably optimizes over-parameterized neural networks”
Simon Du, Xiyu Zhai, Barnabas Poczos and Aarti Singh · 2018
Earlier work this paper cites.
“Learning one-hidden-layer neural networks with landscape design”
Rong Ge, Jason Lee and Tengyu Ma · 2018
Earlier work this paper cites.
“Characterizing implicit bias in terms of optimization geometry”
Suriya Gunasekar, Jason Lee, Daniel Soudry and Nathan Srebro · 2018
Earlier work this paper cites.
“Implicit bias of gradient descent on linear convolutional networks”
Suriya Gunasekar, Jason Lee, Daniel Soudry and Nathan Srebro · 2018
Earlier work this paper cites.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Earlier work this paper cites.
“Learning overparameterized neural networks via stochastic gradient descent on structured data”
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
“Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations”
Yuanzhi Li, Tengyu Ma and Hongyang Zhang · 2018
Earlier work this paper cites.
“Gaussian process behaviour in wide deep neural networks”
Alexander Matthews, Jiri Hron, Mark Rowland, Richard Turner and Zoubin Ghahramani · 2018
Earlier work this paper cites.
“Gaussian process behaviour in wide deep neural networks”
Alexander Matthews, Mark Rowland, Jiri Hron, Richard Turner and Zoubin Ghahramani · 2018
Earlier work this paper cites.
“Convergence results for neural networks via electrodynamics”
Rina Panigrahy, Sushant Sachdeva and Qiuyi Zhang · 2018
Earlier work this paper cites.
“Searching for activation functions”
Prajit Ramachandran, Barret Zoph and Quoc Le · 2018
Earlier work this paper cites.
“Spurious local minima are common in two-layer ReLU neural networks”
Itay Safran and Ohad Shamir · 2018
Earlier work this paper cites.
“The implicit bias of gradient descent on separable data”
Daniel Soudry, Elad Hoffer, Mor Nacson, Suriya Gunasekar and Nathan Srebro · 2018
Cited alongside, same era.
“High-dimensional probability: An introduction with applications in data science”
Roman Vershynin · 2018
Cited alongside, same era.
“A convergence theory for deep learning via over-parameterization”
Zeyuan Allen-Zhu, Yuanzhi Li and Zhao Song · 2019
Cited alongside, same era.
“Implicit regularization in deep matrix factorization”
Sanjeev Arora, Nadav Cohen, Wei Hu and Yuping Luo · 2019
Cited alongside, same era.
“Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks”
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li and Ruosong Wang · 2019
Cited alongside, same era.
“Reconciling modern machine-learning practice and the classical bias–variance trade-off”
“When does gradient descent with logistic loss find interpolating two-layer networks?”
Niladri Chatterji, Philip Long and Peter Bartlett · 2020
Later among the works it cites.
“A generalized neural tangent kernel analysis for two-layer neural networks”
Zixiang Chen, Yuan Cao, Quanquan Gu and Tong Zhang · 2020
Later among the works it cites.
“Analysis of gradient descent on wide two-layer ReLU neural networks” Talk at MSRI, 2020
Lénaïc Chizat · 2020
Later among the works it cites.
“Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss”
Lénaïc Chizat and Francis Bach · 2020
Later among the works it cites.
“Neural Networks learning and memorization with (almost) no over-parameterization”
Amit Daniely · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mikhail Belkin, Daniel Hsu, Siyuan Ma and Soumik Mandal · 2019
Cited alongside, same era.
“Why do larger models generalize better? A theoretical perspective via the XOR problem”
Alon Brutzkus and Amir Globerson · 2019
Cited alongside, same era.
“On lazy training in differentiable programming”
Lénaïc Chizat, Edouard Oyallon and Francis Bach · 2019
Cited alongside, same era.
“Gradient descent finds global minima of deep neural networks”
Simon Du, Jason Lee, Haochuan Li, Liwei Wang and Xiyu Zhai · 2019
Cited alongside, same era.
“Surprises in high-dimensional ridgeless least squares interpolation”
Trevor Hastie, Andrea Montanari, Saharon Rosset and Ryan Tibshirani · 2019
Cited alongside, same era.
“Gradient descent aligns the layers of deep linear networks”
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
“Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks”
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
“Learning parities with neural networks”
Amit Daniely and Eran Malach · 2020
Later among the works it cites.
“Directional convergence and alignment in deep learning”
Ziwei Ji and Matus Telgarsky · 2020
Later among the works it cites.
“Just interpolate: Kernel “ridgeless" regression can generalize”
Tengyuan Liang and Alexander Rakhlin · 2020
Later among the works it cites.
“On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels”
Tengyuan Liang, Alexander Rakhlin and Xiyu Zhai · 2020
Later among the works it cites.
Tengyuan Liang and Pragya Sur · 2020
Later among the works it cites.
“Gradient descent maximizes the margin of homogeneous neural networks”
Kaifeng Lyu and Jian Li · 2020
Later among the works it cites.
“Classification vs regression in overparameterized regimes: Does the loss function matter?”
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu and Anant Sahai · 2020
Later among the works it cites.
“Harmless interpolation of noisy data in regression”
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian and Anant Sahai · 2020
Later among the works it cites.
“Towards moderate overparameterization: global convergence guarantees for training shallow neural networks”
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
“Optimizing mode connectivity via neuron alignment”
Norman Tatro, Pin-Yu Chen, Payel Das, Igor Melnyk, Prasanna Sattigeri and Rongjie Lai · 2020
Later among the works it cites.
“Benign overfitting in ridge regression”
Alexander Tsigler and Peter Bartlett · 2020
Later among the works it cites.
“Gradient descent optimizes over-parameterized deep ReLU networks”
Difan Zou, Yuan Cao, Dongruo Zhou and Quanquan Gu · 2020
Later among the works it cites.
“A deep conditioning treatment of neural networks”
Naman Agarwal, Pranjal Awasthi and Satyen Kale · 2021
Closest in time.
“Finite-sample analysis of interpolating linear classifiers in the overparameterized regime”
Niladri Chatterji and Philip Long · 2021
Closest in time.
“How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?”
Zixiang Chen, Yuan Cao, Difan Zou and Quanquan Gu · 2021
Closest in time.
“On the proliferation of support vectors in high dimensions”
Daniel Hsu, Vidya Muthukumar and Ji Xu · 2021
Closest in time.
“The generalization error of random features regression: Precise asymptotics and the double descent curve”
Song Mei and Andrea Montanari · 2021
Closest in time.
“An Improved Analysis of Training Over-parameterized Deep Neural Networks”
Difan Zou and Quanquan Gu · 2062
Closest in time.