Fetching the paper…
Reading the bibliography…
As part of the effort to understand implicit bias of gradient descent in overparametrized models, several results have shown how the training trajectory on the overparametrized model can be understood as mirror descent on a different objective.
Yang, G · 1902
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K · 1906
Earlier work this paper cites.
Li, Y · 1907
Earlier work this paper cites.
The imbedding problem for riemannian manifolds
Nash, J · 1956
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
Bregman, L. M · 1967
Earlier work this paper cites.
Orbits of families of vector fields and integrability of distributions
Sussmann, H. J · 1973
Earlier work this paper cites.
A relationship between the second derivatives of a convex function and of its conjugate
Crouzeix, J.-P · 1977
Earlier work this paper cites.
An iterative row-action method for interval convex programming
Censor, Y · 1981
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovskij, A. S · 1983
Earlier work this paper cites.
Regularity of the distance function
Foote, R. L · 1984
Earlier work this paper cites.
Isometric embeddings of riemannian manifolds, kyoto, 1990
Gunther, M · 1990
Earlier work this paper cites.
Legendre functions and the method of random bregman projections
Bauschke, H. H · 1997
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A · 2003
Earlier work this paper cites.
Hessian riemannian gradient flows in convex programming
Alvarez, F · 2004
Earlier work this paper cites.
Shape matters: Understanding the implicit bias of the noise covariance
HaoChen, J. Z · 2006
Earlier work this paper cites.
Introduction to differentiable manifolds
Lang, S · 2006
Earlier work this paper cites.
A unifying view on implicit bias in training linear neural networks
Yun, C · 2010
Earlier work this paper cites.
Introduction to Smooth Manifolds
Lee, J. M · 2013
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Earlier work this paper cites.
Convex analysis
Rockafellar, R. T · 2015
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S · 2018
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Cited alongside, same era.
Implicit regularization in deep learning may not be explainable by norms
Razin, N · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Woodworth, B · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Zou, D · 2020
Later among the works it cites.
On the implicit bias of initialization shape: Beyond infinitesimal mirror descent
Azulay, S · 2021
Later among the works it cites.
Label noise sgd provably prefers flat global minimizers
Damian, A · 2021
Later among the works it cites.
Understanding deflation process in over-parametrized tensor decomposition
Ge, R · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ji, Z · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S · 2019
Cited alongside, same era.
The implicit bias of gradient descent on nonseparable data
Ji, Z · 2019
Cited alongside, same era.
Convergence of gradient descent on separable data
Nacson, M. S · 2019
Cited alongside, same era.
The implicit bias of adagrad on separable data
Qian, Q · 2019
Cited alongside, same era.
Mirrorless mirror descent: A natural derivation of mirror descent
Gunasekar, S · 2021
Later among the works it cites.
Deep linear networks dynamics: Low-rank biases induced by initialization scale and l2 regularization
Jacot, A · 2021
Later among the works it cites.
Fast margin maximization via dual acceleration
Ji, Z · 2021
Later among the works it cites.
Characterizing the implicit bias via a primal-dual analysis
Ji, Z · 2021
Later among the works it cites.
Gradient descent on two-layer nets: Margin maximization and simplicity bias
Lyu, K · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Pesme, S · 2021
Later among the works it cites.
Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
Stöger, D · 2021
Later among the works it cites.
Tensor programs iv: Feature learning in infinite-width neural networks
Yang, G · 2021
Later among the works it cites.
The benefits of implicit regularization from sgd in least squares problems
Zou, D · 2021
Later among the works it cites.
Non-convex online learning via algorithmic equivalence
Ghai, U · 2022
Closest in time.
What happens after SGD reaches zero loss? –a mathematical framework
Li, Z · 2022
Closest in time.
Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks
Razin, N · 2022
Closest in time.