Fetching the paper…
Reading the bibliography…
The implicit biases of gradient-based optimization algorithms are conjectured to be a major factor in the success of modern deep learning.
“Adaptive Estimation of a Quadratic Functional by Model Selection”
B. Laurent and P. Massart · 2000
Earlier work this paper cites.
“Sampling from large matrices: An approach through geometric functional analysis”
Mark Rudelson and Roman Vershynin · 2007
Earlier work this paper cites.
“Nonsmooth analysis and control theory”
Francis Clarke, Yuri Ledyaev, Ronald Stern and Peter Wolenski · 2008
Earlier work this paper cites.
“Introduction to the non-asymptotic analysis of random matrices”
Roman Vershynin · 2010
Earlier work this paper cites.
“Approximate KKT points and a proximity measure for termination”
Joydeep Dutta, Kalyanmoy Deb, Rupesh Tulshyan and Ramnik Arora · 2013
Earlier work this paper cites.
“Hanson-wright inequality and sub-gaussian concentration”
Mark Rudelson and Roman Vershynin · 2013
Earlier work this paper cites.
“Sgd learns over-parameterized networks that provably generalize on linearly separable data”
Alon Brutzkus, Amir Globerson, Eran Malach and Shai Shalev-Shwartz · 2018
Earlier work this paper cites.
“Neural Tangent Kernel: Convergence and Generalization in Neural Networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Earlier work this paper cites.
“A Convergence Theory for Deep Learning via Over-Parameterization”
Zeyuan Allen-Zhu, Yuanzhi Li and Zhao Song · 2019
Earlier work this paper cites.
“Implicit regularization in deep matrix factorization”
Sanjeev Arora, Nadav Cohen, Wei Hu and Yuping Luo · 2019
Earlier work this paper cites.
“On exact computation with an infinitely wide neural net”
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov and Ruosong Wang · 2019
Earlier work this paper cites.
“Gradient Descent Provably Optimizes Over-parameterized Neural Networks”
Simon. Du, Xiyu Zhai, Barnabás Póczos and Aarti Singh · 2019
Earlier work this paper cites.
“Algorithm-Dependent Generalization Bounds for Overparameterized Deep Residual Networks”
Spencer Frei, Yuan Cao and Quanquan Gu · 2019
Earlier work this paper cites.
“Gradient descent aligns the layers of deep linear networks”
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
“Theoretical Insights Into the Optimization Landscape of Over-Parameterized Shallow Neural Networks”
Mahdi Soltanolkotabi, Adel Javanmard and Jason. Lee · 2019
Cited alongside, same era.
“Gradient descent optimizes over-parameterized deep ReLU networks”
Difan Zou, Yuan Cao, Dongruo Zhou and Quanquan Gu · 2019
Cited alongside, same era.
“Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss”
Lenaic Chizat and Francis Bach · 2020
Cited alongside, same era.
“The surprising simplicity of the early-time learning dynamics of neural networks”
Wei Hu, Lechao Xiao, Ben Adlam and Jeffrey Pennington · 2020
Cited alongside, same era.
“Gradient descent on two-layer nets: Margin maximization and simplicity bias”
Kaifeng Lyu, Zhiyuan Li, Runzhe Wang and Sanjeev Arora · 2021
Later among the works it cites.
“Towards understanding learning in neural networks with linear teachers”
Roei Sarussi, Alon Brutzkus and Amir Globerson · 2021
Later among the works it cites.
“On Margin Maximization in Linear and ReLU Networks”
Gal Vardi, Ohad Shamir and Nathan Srebro · 2021
Later among the works it cites.
“Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs”
Etienne Boursier, Loucas Pillaud-Vivien and Nicolas Flammarion · 2022
Closest in time.
“Benign Overfitting in Two-layer Convolutional Neural Networks”
Yuan Cao, Zixiang Chen, Mikhail Belkin and Quanquan Gu · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Directional convergence and alignment in deep learning”
Ziwei Ji and Matus Telgarsky · 2020
Cited alongside, same era.
“Gradient descent maximizes the margin of homogeneous neural networks”
Kaifeng Lyu and Jian Li · 2020
Cited alongside, same era.
“The inductive bias of ReLU networks on orthogonally separable data”
Mary Phuong and Christoph Lampert · 2020
Cited alongside, same era.
“Implicit Regularization in Deep Learning May Not Be Explainable by Norms”
Noam Razin and Nadav Cohen · 2020
Cited alongside, same era.
“Finite-sample analysis of interpolating linear classifiers in the overparameterized regime”
Niladri. Chatterji and Philip. Long · 2021
Cited alongside, same era.
“Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise”
Spencer Frei, Yuan Cao and Quanquan Gu · 2021
Cited alongside, same era.
“Proxy Convexity: A Unified Framework for the Analysis of Neural Networks Trained by Gradient Descent”
Spencer Frei and Quanquan Gu · 2021
Cited alongside, same era.
Closest in time.
“Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data”
Spencer Frei, Niladri. Chatterji and Peter. Bartlett · 2022
Closest in time.
“Random Feature Amplification: Feature Learning and Generalization in Neural Networks”
Spencer Frei, Niladri. Chatterji and Peter. Bartlett · 2022
Closest in time.
Itay Safran, Gal Vardi and Jason Lee · 2022
Closest in time.
“Data Augmentation as Feature Manipulation: a story of desert cows and grass cows”
Ruoqi Shen, Sébastien Bubeck and Suriya Gunasekar · 2022
Closest in time.
“Implicit regularization towards rank minimization in relu networks”
Nadav Timor, Gal Vardi and Ohad Shamir · 2022
Closest in time.
“On the Implicit Bias in Deep-Learning Algorithms”
Gal Vardi · 2022
Closest in time.
“Gradient Methods Provably Converge to Non-Robust Networks”
Gal Vardi, Gilad Yehudai and Ohad Shamir · 2022
Closest in time.