Fetching the paper…
Reading the bibliography…
We study the training of finite-width two-layer smoothed ReLU networks for binary classification using the logistic loss.
“Learning polynomials with neural networks”
Alexandr Andoni, Rina Panigrahy, Gregory Valiant and Li Zhang · 1916
Earlier work this paper cites.
“Adaptive estimation of a quadratic functional by model selection”
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
“Introduction to algorithms”
Thomas Cormen, Charles Leiserson, Ronald Rivest and Clifford Stein · 2009
Earlier work this paper cites.
“Convex optimization: algorithms and complexity”
Sébastien Bubeck · 2015
Earlier work this paper cites.
“In search of the real inductive bias: On the role of implicit regularization in deep learning.”
Behnam Neyshabur, Ryota Tomioka and Nathan Srebro · 2015
Earlier work this paper cites.
“Convergence analysis of two-layer neural networks with ReLU activation”
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht and Oriol Vinyals · 2017
Earlier work this paper cites.
“On the learnability of fully-connected neural networks”
Yuchen Zhang, Jason Lee, Martin Wainwright and Michael Jordan · 2017
Earlier work this paper cites.
“Recovery Guarantees for One-hidden-layer Neural Networks”
Kai Zhong, Zhao Song, Prateek Jain, Peter Bartlett and Inderjit. Dhillon · 2017
Earlier work this paper cites.
“Overfitting or perfect fitting? Risk bounds for classification and regression rules that interpolate”
Mikhail Belkin, Daniel Hsu and Partha Mitra · 2018
Earlier work this paper cites.
“SGD learns over-parameterized networks that provably generalize on linearly separable data”
Alon Brutzkus, Amir Globerson, Eran Malach and Shai Shalev-Shwartz · 2018
Earlier work this paper cites.
“On the global convergence of gradient descent for over-parameterized models using optimal transport”
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
“Gradient descent provably optimizes over-parameterized neural networks”
Simon Du, Xiyu Zhai, Barnabas Poczos and Aarti Singh · 2018
Earlier work this paper cites.
“Learning one-hidden-layer neural networks with landscape design”
Rong Ge, Jason Lee and Tengyu Ma · 2018
Earlier work this paper cites.
“Characterizing implicit bias in terms of optimization geometry”
Suriya Gunasekar, Jason Lee, Daniel Soudry and Nathan Srebro · 2018
Earlier work this paper cites.
“Implicit bias of gradient descent on linear convolutional networks”
Suriya Gunasekar, Jason Lee, Daniel Soudry and Nathan Srebro · 2018
Earlier work this paper cites.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Earlier work this paper cites.
“Learning overparameterized neural networks via stochastic gradient descent on structured data”
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
“Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations”
Yuanzhi Li, Tengyu Ma and Hongyang Zhang · 2018
Earlier work this paper cites.
“Convergence results for neural networks via electrodynamics”
Rina Panigrahy, Sushant Sachdeva and Qiuyi Zhang · 2018
Earlier work this paper cites.
“Searching for activation functions”
Prajit Ramachandran, Barret Zoph and Quoc Le · 2018
Earlier work this paper cites.
“Spurious local minima are common in two-layer ReLU neural networks”
Itay Safran and Ohad Shamir · 2018
Earlier work this paper cites.
“The implicit bias of gradient descent on separable data”
Daniel Soudry, Elad Hoffer, Mor Nacson, Suriya Gunasekar and Nathan Srebro · 2018
Cited alongside, same era.
“High-dimensional probability: An introduction with applications in data science”
Roman Vershynin · 2018
Cited alongside, same era.
“A convergence theory for deep learning via over-parameterization”
Zeyuan Allen-Zhu, Yuanzhi Li and Zhao Song · 2019
Cited alongside, same era.
“Implicit regularization in deep matrix factorization”
Sanjeev Arora, Nadav Cohen, Wei Hu and Yuping Luo · 2019
Cited alongside, same era.
“Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks”
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li and Ruosong Wang · 2019
Cited alongside, same era.
“Reconciling modern machine-learning practice and the classical bias–variance trade-off”
“A corrective view of neural networks: Representation, memorization and learning”
Guy Bresler and Dheeraj Nagaraj · 2020
Closest in time.
“When does gradient descent with logistic loss find interpolating two-layer networks?”
Niladri Chatterji, Philip Long and Peter Bartlett · 2020
Closest in time.
“A generalized neural tangent kernel analysis for two-layer neural networks”
Zixiang Chen, Yuan Cao, Quanquan Gu and Tong Zhang · 2020
Closest in time.
“Analysis of gradient descent on wide two-layer ReLU neural networks” Talk at MSRI, 2020
Lénaïc Chizat · 2020
Closest in time.
“Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss”
Lénaïc Chizat and Francis Bach · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mikhail Belkin, Daniel Hsu, Siyuan Ma and Soumik Mandal · 2019
Cited alongside, same era.
“Why do larger models generalize better? A theoretical perspective via the XOR problem”
Alon Brutzkus and Amir Globerson · 2019
Cited alongside, same era.
“On lazy training in differentiable programming”
Lénaïc Chizat, Edouard Oyallon and Francis Bach · 2019
Cited alongside, same era.
“Gradient descent finds global minima of deep neural networks”
Simon Du, Jason Lee, Haochuan Li, Liwei Wang and Xiyu Zhai · 2019
Cited alongside, same era.
“Surprises in high-dimensional ridgeless least squares interpolation”
Trevor Hastie, Andrea Montanari, Saharon Rosset and Ryan Tibshirani · 2019
Cited alongside, same era.
“Gradient descent aligns the layers of deep linear networks”
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
“Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks”
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Amit Daniely · 2020
Closest in time.
“Learning parities with neural networks”
Amit Daniely and Eran Malach · 2020
Closest in time.
“On the proliferation of support vectors in high dimensions”
Daniel Hsu, Vidya Muthukumar and Ji Xu · 2020
Closest in time.
“Directional convergence and alignment in deep learning”
Ziwei Ji and Matus Telgarsky · 2020
Closest in time.
“Just interpolate: Kernel “ridgeless" regression can generalize”
Tengyuan Liang and Alexander Rakhlin · 2020
Closest in time.
“On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels”
Tengyuan Liang, Alexander Rakhlin and Xiyu Zhai · 2020
Closest in time.
Tengyuan Liang and Pragya Sur · 2020
Closest in time.
“Gradient descent maximizes the margin of homogeneous neural networks”
Kaifeng Lyu and Jian Li · 2020
Closest in time.
“Classification vs regression in overparameterized regimes: Does the loss function matter?”
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu and Anant Sahai · 2020
Closest in time.
“Harmless interpolation of noisy data in regression”
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian and Anant Sahai · 2020
Closest in time.
“Towards moderate overparameterization: global convergence guarantees for training shallow neural networks”
Samet Oymak and Mahdi Soltanolkotabi · 2020
Closest in time.
“Optimizing mode connectivity via neuron alignment”
Norman Tatro, Pin-Yu Chen, Payel Das, Igor Melnyk, Prasanna Sattigeri and Rongjie Lai · 2020
Closest in time.
“Benign overfitting in ridge regression”
Alexander Tsigler and Peter Bartlett · 2020
Closest in time.
“Gradient descent optimizes over-parameterized deep ReLU networks”
Difan Zou, Yuan Cao, Dongruo Zhou and Quanquan Gu · 2020
Closest in time.
“Finite-sample analysis of interpolating linear classifiers in the overparameterized regime”
Niladri Chatterji and Philip Long · 2021
Closest in time.
“When does gradient descent with logistic loss interpolate using deep networks with smoothed ReLU activations?”
Niladri Chatterji, Philip Long and Peter Bartlett · 2021
Closest in time.