Fetching the paper…
Reading the bibliography…
Matrix factorization is a simple and natural test-bed to investigate the implicit regularization of gradient descent.
A refined primal-dual analysis of the implicit bias
Ziwei Ji and Matus Telgarsky · 1906
Earlier work this paper cites.
Ensembles semi-analytiques
Stanisław Łojasiewicz · 1965
Earlier work this paper cites.
Generalized gradients and applications
Frank H. Clarke · 1975
Earlier work this paper cites.
Optimization and Nonsmooth Analysis
Frank H Clarke · 1990
Earlier work this paper cites.
The clarke and michel-penot subdifferentials of the eigenvalues of a symmetric matrix
Jean-Baptiste Hiriart-Urruty and A. S. Lewis · 1999
Earlier work this paper cites.
Nonsmooth analysis and control theory , volume 178
Francis H. Clarke, Yuri S. Ledyaev, Ronald J. Stern, and Peter R. Wolenski · 2008
Earlier work this paper cites.
On the equivalence of weak learnability and linear separability: New relaxations and efficient boosting algorithms
Shai Shalev-Shwartz and Yoram Singer · 2010
Earlier work this paper cites.
Julia: A fast dynamic language for technical computing
Jeff Bezanson, Stefan Karpinski, Viral B Shah, and Alan Edelman · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Differential equations and dynamical systems , volume 7
Lawrence Perko · 2013
Earlier work this paper cites.
Rank-one matrix pursuit for matrix completion
Zheng Wang, Ming-Jun Lai, Zhaosong Lu, Wei Fan, Hasan Davulcu, and Jieping Ye · 2014
Earlier work this paper cites.
CVXPY: A Python-embedded modeling language for convex optimization
Steven Diamond and Stephen Boyd · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
Greedy learning of generalized low-rank models
Quanming Yao and James Tin Yau Kwok · 2016
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
On approximation guarantees for greedy low rank optimization
Rajiv Khanna, Ethan Elenberg, Alexandros G Dimakis, and Sahand Negahban · 2017
Cited alongside, same era.
First-order methods almost always avoid saddle points
Jason D Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2017
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Later among the works it cites.
Nonconvex optimization meets low-rank matrix factorization: An overview
Yuejie Chi, Yue M Lu, and Yuxin Chen · 2019
Later among the works it cites.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Later among the works it cites.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Later among the works it cites.
Structured Low-Rank Matrix Factorization: Global Optimality, Algorithms, and Applications
Benjamin D. Haeffele and René Vidal · 2019
Later among the works it cites.
First-order methods almost always avoid saddle points: The case of vanishing step-sizes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
A rewriting system for convex optimization problems
Akshay Agrawal, Robin Verschueren, Steven Diamond, and Stephen Boyd · 2018
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, N Cohen, and Elad Hazan · 2018
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
Simon Du and Jason Lee · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Ioannis Panageas, Georgios Piliouras, and Xiao Wang · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
On implicit regularization: Morse functions and applications to matrix factorization
Mohamed Ali Belabbas · 2020
Closest in time.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lénaïc Chizat and Francis Bach · 2020
Closest in time.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2020
Closest in time.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Closest in time.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Closest in time.