Fetching the paper…
Reading the bibliography…
We study the implicit bias of generic optimization methods, such as mirror descent, natural gradient descent, and steepest descent with respect to different potentials and norms, when optimizing underdetermined linear regression or separable linear classification problems.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
L. M. Bregman · 1967
Earlier work this paper cites.
Elementary analysis
Kenneth A Ross · 1980
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. Nemirovskii and D. Yudin · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o (1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Exponentiated gradient versus gradient descent for linear predictors
Jyrki Kivinen and Manfred K Warmuth · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
S. I. Amari · 1998
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Jerome H Friedman · 2001
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2003
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
A. Beck and M. Teboulle · 2003
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Least angle regression
B. Efron, T. Hastie, I. Johnstone, and R. Tibshirani · 2004
Cited alongside, same era.
Boosting as a regularized path to a maximum margin classifier
S. Rosset, J. Zhu, and T. Hastie · 2004
Cited alongside, same era.
The dynamics of adaboost: Cyclic behavior and convergence of margins
Cynthia Rudin, Ingrid Daubechies, and Robert E Schapire · 2004
Cited alongside, same era.
Generalization error bounds for collaborative prediction with low-rank matrices
Nathan Srebro, Noga Alon, and Tommi S Jaakkola · 2005
Cited alongside, same era.
Boosting with early stopping: Convergence and consistency
Tong Zhang, Bin Yu, et al · 2005
Cited alongside, same era.
Exact matrix completion via convex optimization
E. J. Candes and B. Recht · 2009
Cited alongside, same era.
Margins, shrinkage and boosting
Matus Telgarsky · 2013
Later among the works it cites.
Adam: A method for stochastic optimisation
D Kingma and Jimmy Ba Adam · 2015
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Later among the works it cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Later among the works it cites.
Elad Hoffer, I Hubara, and D. Soudry · 2017
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix recovery
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A concrete approach to classical analysis , volume 14
Marian Muresan and Marian Muresan · 2009
Cited alongside, same era.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
B. Recht, M. Fazel, and P. A. Parrilo · 2010
Cited alongside, same era.
On the equivalence of weak learnability and linear separability: New relaxations and efficient boosting algorithms
Shai Shalev-Shwartz and Yoram Singer · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Boosting: Foundations and algorithms
Robert E Schapire and Yoav Freund · 2012
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan R Salakhutdinov, and Nati Srebro
Cited in the paper.
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2017
Later among the works it cites.
Geometry of optimization and implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Later among the works it cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, and Nathan Srebro · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Don’t Decay the Learning Rate, Increase the Batch Size
Le Smith, Kindermans · 2018
Closest in time.