Fetching the paper…
Reading the bibliography…
Linear classifiers and leaky ReLU networks trained by gradient flow on the logistic loss have an implicit bias towards solutions which satisfy the Karush--Kuhn--Tucker (KKT) conditions for margin maximization.
“The Geometry of Logconcave Functions and Sampling Algorithms”
László Lovász and Santosh Vempala · 2007
Earlier work this paper cites.
“Nonsmooth analysis and control theory”
Francis Clarke, Yuri Ledyaev, Ronald Stern and Peter Wolenski · 2008
Earlier work this paper cites.
“Approximate KKT points and a proximity measure for termination”
Joydeep Dutta, Kalyanmoy Deb, Rupesh Tulshyan and Ramnik Arora · 2013
Earlier work this paper cites.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht and Oriol Vinyals · 2017
Earlier work this paper cites.
“Overfitting or perfect fitting? Risk bounds for classification and regression rules that interpolate”
Mikhail Belkin, Daniel Hsu and Partha Mitra · 2018
Earlier work this paper cites.
“Characterizing implicit bias in terms of optimization geometry”
Suriya Gunasekar, Jason Lee, Daniel Soudry and Nathan Srebro · 2018
Earlier work this paper cites.
“Risk and parameter convergence of logistic regression”
Ziwei Ji and Matus Telgarsky · 2018
Earlier work this paper cites.
“The Implicit Bias of Gradient Descent on Separable Data”
Daniel Soudry, Elad Hoffer, Mor Nacson, Suriya Gunasekar and Nathan Srebro · 2018
Earlier work this paper cites.
“High-Dimensional Probability: An Introduction with Applications in Data Science”
Roman Vershynin · 2018
Earlier work this paper cites.
“The generalization error of random features regression: Precise asymptotics and the double descent curve”
Song Mei and Andrea Montanari · 2019
Earlier work this paper cites.
“Convergence of gradient descent on separable data”
Mor Nacson, Jason Lee, Suriya Gunasekar, Pedro Savarese, Nathan Srebro and Daniel Soudry · 2019
Earlier work this paper cites.
“Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate”
Mor Nacson, Nathan Srebro and Daniel Soudry · 2019
Earlier work this paper cites.
“Consistency of interpolation with Laplace kernels is a high-dimensional phenomenon”
Alexander Rakhlin and Xiyu Zhai · 2019
Earlier work this paper cites.
“Benign Overfitting in Linear Regression”
Peter. Bartlett, Philip. Long, Gábor Lugosi and Alexander Tsigler · 2020
Earlier work this paper cites.
“Two models of double descent for weak features”
Mikhail Belkin, Daniel Hsu and Ji Xu · 2020
Earlier work this paper cites.
“On the robustness of the minimum ℓ 2 \ell_{2} interpolator”
Geoffrey Chinot and Matthieu Lerasle · 2020
Earlier work this paper cites.
“Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss”
Lenaic Chizat and Francis Bach · 2020
Earlier work this paper cites.
“Surprises in High-Dimensional Ridgeless Least Squares Interpolation”
Trevor Hastie, Andrea Montanari, Saharon Rosset and Ryan. Tibshirani · 2020
Earlier work this paper cites.
“Gradient descent follows the regularization path for general losses”
Ziwei Ji, Miroslav Dudik, Robert Schapire and Matus Telgarsky · 2020
Earlier work this paper cites.
“Directional convergence and alignment in deep learning”
Ziwei Ji and Matus Telgarsky · 2020
Earlier work this paper cites.
“Just interpolate: Kernel “ridgeless” regression can generalize”
Tengyuan Liang and Alexander Rakhlin · 2020
Earlier work this paper cites.
“On the Multiple Descent of Minimum-Norm Interpolants and Restricted Lower Isometry of Kernels”
Tengyuan Liang, Alexander Rakhlin and Xiyu Zhai · 2020
Cited alongside, same era.
“Gradient Descent Maximizes the Margin of Homogeneous Neural Networks”
Kaifeng Lyu and Jian Li · 2020
Cited alongside, same era.
Andrea Montanari, Feng Ruan, Youngtak Sohn and Jun Yan · 2020
Cited alongside, same era.
“Harmless interpolation of noisy data in regression”
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian and Anant Sahai · 2020
Cited alongside, same era.
“In defense of uniform convergence: Generalization via derandomization with an application to interpolating predictors”
Jeffrey Negrea, Gintare Dziugaite and Daniel Roy · 2020
Cited alongside, same era.
“Gradient descent on two-layer nets: Margin maximization and simplicity bias”
Kaifeng Lyu, Zhiyuan Li, Runzhe Wang and Sanjeev Arora · 2021
Later among the works it cites.
“Classification vs regression in overparameterized regimes: Does the loss function matter?”
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu and Anant Sahai · 2021
Later among the works it cites.
“Towards understanding learning in neural networks with linear teachers”
Roei Sarussi, Alon Brutzkus and Amir Globerson · 2021
Later among the works it cites.
“On Margin Maximization in Linear and ReLU Networks”
Gal Vardi, Ohad Shamir and Nathan Srebro · 2021
Later among the works it cites.
“Benign Overfitting in Multiclass Classification: All Roads Lead to Interpolation”
Ke Wang, Vidya Muthukumar and Christos Thrampoulidis · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“The inductive bias of ReLU networks on orthogonally separable data”
Mary Phuong and Christoph Lampert · 2020
Cited alongside, same era.
“Gradient methods never overfit on separable data”
Ohad Shamir · 2020
Cited alongside, same era.
“Theoretical insights into multiclass classification: A high-dimensional asymptotic view”
Christos Thrampoulidis, Samet Oymak and Mahdi Soltanolkotabi · 2020
Cited alongside, same era.
“Benign overfitting in ridge regression”
A. Tsigler and P.. Bartlett · 2020
Cited alongside, same era.
“On the Optimal Weighted ℓ 2 \ell_{2} Regularization in Overparameterized Linear Regression”
Denny Wu and Ji Xu · 2020
Cited alongside, same era.
“Failures of model-dependent generalization bounds for least-norm interpolation”
Peter Bartlett and Philip Long · 2021
Cited alongside, same era.
“Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian Mixtures”
Yuan Cao, Quanquan Gu and Mikhail Belkin · 2021
Cited alongside, same era.
Ke Wang and Christos Thrampoulidis · 2021
Later among the works it cites.
“Benign overfitting in two-layer convolutional neural networks”
Yuan Cao, Zixiang Chen, Mikhail Belkin and Quanquan Gu · 2022
Later among the works it cites.
“Fast rates for noisy interpolation require rethinking the effect of inductive bias”
Konstantin Donhauser, Nicolo Ruggeri, Stefan Stojanovic and Fanny Yang · 2022
Later among the works it cites.
“Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data”
Spencer Frei, Niladri. Chatterji and Peter. Bartlett · 2022
Later among the works it cites.
“The Asymmetric Maximum Margin Bias of Quasi-Homogeneous Neural Networks”
Daniel Kunin, Atsushi Yamamura, Chao Ma and Surya Ganguli · 2022
Later among the works it cites.
“Benign, tempered, or catastrophic: A taxonomy of overfitting”
Neil Mallinar, James Simon, Amirhesam Abedsoltan, Parthe Pandit, Mikhail Belkin and Preetum Nakkiran · 2022
Later among the works it cites.
“Harmless interpolation in regression and classification with structured features”
Andrew McRae, Santhosh Karnik, Mark Davenport and Vidya Muthukumar · 2022
Later among the works it cites.
Itay Safran, Gal Vardi and Jason Lee · 2022
Later among the works it cites.
“The implicit bias of benign overfitting”
Ohad Shamir · 2022
Later among the works it cites.
“Implicit regularization towards rank minimization in relu networks”
Nadav Timor, Gal Vardi and Ohad Shamir · 2022
Later among the works it cites.
“On the Implicit Bias in Deep-Learning Algorithms”
Gal Vardi · 2022
Later among the works it cites.
“Gradient Methods Provably Converge to Non-Robust Networks”
Gal Vardi, Gilad Yehudai and Ohad Shamir · 2022
Later among the works it cites.
“Tight bounds for minimum l1-norm interpolation of noisy data”
Guillaume Wang, Konstantin Donhauser and Fanny Yang · 2022
Later among the works it cites.
“A Non-Asymptotic Moreau Envelope Theory for High-Dimensional Generalized Linear Models”
Lijia Zhou, Frederic Koehler, Pragya Sur, Danica Sutherland and Nathan Srebro · 2022
Later among the works it cites.
“Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data”
Spencer Frei, Gal Vardi, Peter. Bartlett, Nathan Srebro and Wei Hu · 2023
Closest in time.