Fetching the paper…
Reading the bibliography…
Diagonal linear networks (DLNs) are a toy simplification of artificial neural networks; they consist in a quadratic reparametrization of linear regression inducing a sparse implicit regularization.
Über die lage der integralkurven gewöhnlicher differentialgleichungen
Mitio Nagumo · 1942
Earlier work this paper cites.
Stability in models of mutualism
Bean-San Goh · 1979
Earlier work this paper cites.
Linear Operators, Part 1: General Theory , volume 10
Nelson Dunford and Jacob Schwartz · 1988
Earlier work this paper cites.
Evolutionary Games and Population Dynamics
Josef Hofbauer and Karl Sigmund · 1998
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Set-Theoretic Methods in Control
Franco Blanchini and Stefano Miani · 2008
Earlier work this paper cites.
The Linear Complementarity Problem
Richard Cottle, Jong-Shi Pang, and Richard Stone · 2009
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2009
Earlier work this paper cites.
Noisy heteroclinic networks
Yuri Bakhtin · 2011
Earlier work this paper cites.
Optimization with sparsity-inducing penalties
Francis Bach, Rodolphe Jenatton, Julien Mairal, and Guillaume Obozinski · 2012
Earlier work this paper cites.
Deep learning
Yann Le Cun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Lotka–Volterra dynamical systems
Stephen Baigent · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Cited alongside, same era.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2019
Cited alongside, same era.
A mathematical theory of semantic development in deep neural networks
Andrew Saxe, James McClelland, and Surya Ganguli · 2019
Cited alongside, same era.
Implicit regularization for optimal sparse recovery
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2019
Shape matters: Understanding the implicit bias of the noise covariance
Jeff HaoChen, Colin Wei, Jason Lee, and Tengyu Ma · 2021
Later among the works it cites.
Arthur Jacot, François Ged, Berfin Şimşek, Clément Hongler, and Franck Gabriel · 2021
Later among the works it cites.
Implicit sparse regularization: The impact of depth and early stopping
Jiangyuan Li, Thanh Nguyen, Chinmay Hegde, and Ka Wai Wong · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Smooth bilevel programming for sparse regularization
Clarice Poon and Gabriel Peyré · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Peng Zhao, Yun Yang, and Qiao-Chu He · 2019
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Cited alongside, same era.
Gradient descent for deep matrix factorization: Dynamics and implicit bias towards low rank
Hung-Hsu Chou, Carsten Gieshoff, Johannes Maly, and Holger Rauhut · 2020
Cited alongside, same era.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Cited alongside, same era.
On the implicit bias of initialization shape: Beyond infinitesimal mirror descent
Shahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake Woodworth, Nathan Srebro, Amir Globerson, and Daniel Soudry · 2021
Cited alongside, same era.
Implicit regularization in tensor factorization
Noam Razin, Asaf Maman, and Nadav Cohen · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
Etienne Boursier, Loucas Pillaud-Viven, and Nicolas Flammarion · 2022
Closest in time.
Implicit regularization with polynomial growth in deep tensor factorization
Kais Hariz, Hachem Kadri, Stéphane Ayache, Maher Moakher, and Thierry Artières · 2022
Closest in time.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Closest in time.
Label noise (stochastic) gradient descent implicitly solves the Lasso for quadratic parametrisation
Loucas Pillaud-Vivien, Julien Reygner, and Nicolas Flammarion · 2022
Closest in time.
Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks
Noam Razin, Asaf Maman, and Nadav Cohen · 2022
Closest in time.
More is less: inducing sparsity via overparameterization
Hung-Hsu Chou, Johannes Maly, and Holger Rauhut · 2023
Closest in time.