Fetching the paper…
Reading the bibliography…
We study the implicit bias of gradient flow (i.e., gradient descent with infinitesimal step size) on linear neural network training.
A refined primal-dual analysis of the implicit bias
Ziwei Ji and Matus Telgarsky · 1906
Earlier work this paper cites.
Singular values and eigenvalues of tensors: a variational approach
Lek-Heng Lim · 2005
Earlier work this paper cites.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2010
Earlier work this paper cites.
Most tensor problems are NP-hard
Christopher J Hillar and Lek-Heng Lim · 2013
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Earlier work this paper cites.
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
Peter Bartlett, Dave Helmbold, and Philip Long · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Width provably matters in optimization for deep linear neural networks
Simon S Du and Wei Hu · 2019
Cited alongside, same era.
Global convergence of gradient descent for deep linear residual networks
Lei Wu, Qingcan Wang, and Chao Ma · 2019
Cited alongside, same era.
Identity crisis: Memorization and generalization under extreme overparameterization
Provable benefit of orthogonal initialization in optimizing deep linear networks
Wei Hu, Lechao Xiao, and Jeffrey Pennington · 2020
Closest in time.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Closest in time.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Closest in time.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Edward Moroshko, Suriya Gunasekar, Blake Woodworth, Jason D Lee, Nathan Srebro, and Daniel Soudry · 2020
Closest in time.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Michael C Mozer, and Yoram Singer · 2019
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Cited alongside, same era.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2020
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu
Cited in the paper.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo
Cited in the paper.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai
Cited in the paper.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh
Cited in the paper.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Closest in time.
Inductive bias of multi-channel linear convolutional networks with bounded weight norm
Meena Jagadeesan, Ilya Razenshteyn, and Suriya Gunasekar · 2021
Closest in time.