Fetching the paper…
Reading the bibliography…
In this paper we fully describe the trajectory of gradient flow over diagonal linear networks in the limit of vanishing initialisation.
Orthogonal matching pursuit: Recursive function approximation with applications to wavelet decomposition
Yagyensh Chandra Pati, Ramin Rezaiifar, and Perinkulam Sambamurthy Krishnaprasad · 1993
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Adaptive greedy approximations
Geoff Davis, Stephane Mallat, and Marco Avellaneda · 1997
Earlier work this paper cites.
Robust heteroclinic cycles
M. Krupa · 1997
Earlier work this paper cites.
On gradients of functions definable in o-minimal structures
Krzysztof Kurdyka · 1998
Earlier work this paper cites.
Heteroclinic networks in coupled cell systems
Peter Ashwin and Michael Field · 1999
Earlier work this paper cites.
A new approach to variable selection in least squares problems
Michael R Osborne, Brett Presnell, and Berwin A Turlach · 2000
Earlier work this paper cites.
Atomic decomposition by basis pursuit
Scott Shaobing Chen, David L Donoho, and Michael A Saunders · 2001
Earlier work this paper cites.
Singular Riemannian barrier methods and gradient-projection dynamical systems for constrained optimization
H. Attouch, J. Bolte, P. Redont, and M. Teboulle · 2004
Earlier work this paper cites.
Least angle regression
Bradley Efron, Trevor Hastie, Iain Johnstone, and Robert Tibshirani · 2004
Earlier work this paper cites.
Evolution of rate-independent systems
Alexander Mielke · 2005
Earlier work this paper cites.
An iterative regularization method for total variation-based image restoration
Stanley Osher, Martin Burger, Donald Goldfarb, Jinjun Xu, and Wotao Yin · 2005
Earlier work this paper cites.
Stable signal recovery from incomplete and inaccurate measurements
E. Candès, J. Romberg, and T. Tao · 2006
Earlier work this paper cites.
On the rate-independent limit of systems with dry friction and small viscosity
Messoud A. Efendiev and Alexander Mielke · 2006
Earlier work this paper cites.
The restricted isometry property and its implications for compressed sensing
Emmanuel J Candes · 2008
Earlier work this paper cites.
Bregman iterative algorithms for l1-minimization with applications to compressed sensing: Siam journal on imaging sciences, 1, 143–168
W Yin, S Osher, D Goldfarb, and J Darbon · 2008
Earlier work this paper cites.
Modeling solutions with jumps for rate-independent systems on metric spaces
Alexander Mielke, Riccarda Rossi, and Giuseppe Savaré · 2009
Earlier work this paper cites.
Split bregman methods and frame based image restoration
Jian-Feng Cai, Stanley Osher, and Zuowei Shen · 2010
Earlier work this paper cites.
Noisy heteroclinic networks
Yuri Bakhtin · 2011
Cited alongside, same era.
Complexity analysis of the lasso regularization path
Julien Mairal and Bin Yu · 2012
Cited alongside, same era.
Variational convergence of gradient flows and rate-independent evolutions in metric spaces
Alexander Mielke, Riccarda Rossi, and Giuseppe Savaré · 2012
Cited alongside, same era.
An adaptive inverse scale space method for compressed sensing
Martin Burger, Michael Möller, Martin Benning, and Stanley Osher · 2013
Cited alongside, same era.
The lasso problem and uniqueness
Ryan J. Tibshirani · 2013
Cited alongside, same era.
A dual split bregman method for fast l1 minimization
Yi Yang, Michael Möller, and Stanley Osher · 2013
Cited alongside, same era.
Exponentiated gradient meets gradient descent
Udaya Ghai, Elad Hazan, and Yoram Singer · 2020
Later among the works it cites.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
On the implicit bias of initialization shape: Beyond infinitesimal mirror descent
Shahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E Woodworth, Nathan Srebro, Amir Globerson, and Daniel Soudry · 2021
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason Lee, and Tengyu Ma · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Global optimality of local search for low rank matrix recovery
Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Implicit regularization in deep learning
Behnam Neyshabur · 2017
Cited alongside, same era.
Non-square matrix sensing without spurious local minima via the Burer-Monteiro approach
Dohyung Park, Anastasios Kyrillidis, Constantine Carmanis, and Sujay Sanghavi · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Arthur Jacot, François Ged, Berfin Şimşek, Clément Hongler, and Franck Gabriel · 2021
Later among the works it cites.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Implicit regularization in tensor factorization
Noam Razin, Asaf Maman, and Nadav Cohen · 2021
Later among the works it cites.
Incremental learning in diagonal linear networks
Raphaël Berthier · 2022
Later among the works it cites.
Gradient flow dynamics of shallow reLU networks for square loss and orthogonal inputs
Etienne Boursier, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Later among the works it cites.
Algorithmic regularization in model-free overparametrized asymmetric matrix factorization
Liwei Jiang, Yudong Chen, and Lijun Ding · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2023
Closest in time.
Transformers learn through gradual rank increase
Enric Boix-Adsera, Etai Littwin, Emmanuel Abbe, Samy Bengio, and Joshua Susskind · 2023
Closest in time.
(s)gd over diagonal linear networks: Implicit regularisation, large stepsizes and edge of stability
Mathieu Even, Scott Pesme, Suriya Gunasekar, and Nicolas Flammarion · 2023
Closest in time.
Understanding incremental learning of gradient descent: A fine-grained analysis of matrix sensing
Jikai Jin, Zhiyuan Li, Kaifeng Lyu, Simon S Du, and Jason D Lee · 2023
Closest in time.
Johan S Wind, Vegard Antun, and Anders C Hansen · 2023
Closest in time.