Fetching the paper…
Reading the bibliography…
Over the past years, there has been significant interest in understanding the implicit bias of gradient descent optimization and its connection to the generalization properties of overparametrized neural networks.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
L.M. Bregman · 1967
Earlier work this paper cites.
Linear least squares and quadratic programming
Gene H. Golub and Michael A. Saunders · 1970
Earlier work this paper cites.
On the numerical solution of constrained least-squares problems
Josef Stoer · 1971
Earlier work this paper cites.
A numerically stable form of the simplex algorithm
Philip E Gill and Walter Murray · 1973
Earlier work this paper cites.
A family of embedded runge-kutta formulae
J.R. Dormand and P.J. Prince · 1980
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Yurii E Nesterov · 1983
Earlier work this paper cites.
Possible generalization of boltzmann-gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
Two-point step size gradient methods
Jonathan Barzilai and Jonathan M Borwein · 1988
Earlier work this paper cites.
On iterative algorithms for linear least squares problems with bound constraints
Michel Bierlaire, Ph L Toint, and Daniel Tuyttens · 1991
Earlier work this paper cites.
Maximum entropy and the nearly black object
David L Donoho, Iain M Johnstone, Jeffrey C Hoch, and Alan S Stern · 1992
Earlier work this paper cites.
Solving least squares problems
Charles L Lawson and Richard J Hanson · 1995
Earlier work this paper cites.
Numerical Methods for Least Squares Problems
Ake Björck · 1996
Earlier work this paper cites.
A fast non-negativity-constrained least squares algorithm
Rasmus Bro and Sijmen De Jong · 1997
Earlier work this paper cites.
The Barzilai and Borwein gradient method for the large scale unconstrained minimization problem
Marcos Raydan · 1997
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The nature of statistical learning theory
Vladimir Vapnik · 1999
Earlier work this paper cites.
Fast algorithm for the solution of large-scale non-negativity-constrained least squares problems
Mark H Van Benthem and Michael R Keenan · 2004
Earlier work this paper cites.
Projected Barzilai-Borwein methods for large-scale box-constrained quadratic programming
Yu-Hong Dai and Roger Fletcher · 2005
Earlier work this paper cites.
Sparse nonnegative solution of underdetermined linear equations by linear programming
David L Donoho and Jared Tanner · 2005
Earlier work this paper cites.
On the Barzilai-Borwein method
Roger Fletcher · 2005
Earlier work this paper cites.
An interior point Newton-like method for non-negative least-squares problems with degenerate solution
Stefania Bellavia, Maria Macconi, and Benedetta Morini · 2006
Earlier work this paper cites.
Near-Optimal Signal Recovery From Random Projections: Universal Encoding Strategies?
Emmanuel J. Candès and Terence Tao · 2006
Earlier work this paper cites.
Compressed sensing
David L Donoho · 2006
Earlier work this paper cites.
Projected gradient methods for nonnegative matrix factorization
Chih-Jen Lin · 2007
Earlier work this paper cites.
On the uniqueness of nonnegative sparse solutions to underdetermined systems of equations
Alfred M Bruckstein, Michael Elad, and Michael Zibulevsky · 2008
Earlier work this paper cites.
Nonnegative least-squares image deblurring: improved gradient projection approaches
Federico Benvenuto, Riccardo Zanella, Luca Zanni, and Mario Bertero · 2009
Earlier work this paper cites.
Nonnegativity constraints in numerical analysis
Donghui Chen and Robert J Plemmons · 2010
Earlier work this paper cites.
Counting the faces of randomly-projected hypercubes and orthants, with applications
David L Donoho and Jared Tanner · 2010
Earlier work this paper cites.
A unique “nonnegative” solution to an underdetermined system: From vectors to matrices
Meng Wang, Weiyu Xu, and Ao Tang · 2010
Earlier work this paper cites.
IsoformEx: isoform level gene expression estimation using weighted non-negative least squares from mRNA-Seq data
Hyunsoo Kim, Yingtao Bi, Sharmistha Pal, Ravi Gupta, and Ramana V Davuluri · 2011
Earlier work this paper cites.
Nonnegative least-mean-square algorithm
Jie Chen, Cédric Richard, José Carlos M Bermudez, and Paul Honeine · 2011
Earlier work this paper cites.
Efficient parallel nonnegative least squares on multicore architectures
Yuancheng Luo and Ramani Duraiswami · 2011
Earlier work this paper cites.
Sparse recovery by thresholded non-negative least squares
Martin Slawski and Matthias Hein · 2011
Earlier work this paper cites.
A method for finding structured sparse solutions to nonnegative least squares problems with applications
Ernie Esser, Yifei Lou, and Jack Xin · 2013
Cited alongside, same era.
A non-monotonic method for large-scale non-negative least squares
Dongmin Kim, Suvrit Sra, and Inderjit S Dhillon · 2013
Cited alongside, same era.
A Mathematical Introduction to Compressive Sensing
Simon Foucart and Holger Rauhut · 2013
Cited alongside, same era.
Non-negative least squares for high-dimensional linear models: Consistency and sparse recovery without regularization
Martin Slawski and Matthias Hein · 2013
Cited alongside, same era.
Sign-constrained least squares estimation for high-dimensional regression
Nicolai Meinshausen · 2013
Cited alongside, same era.
Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward–backward splitting, and regularized gauss–seidel methods
Low-rank regularization and solution uniqueness in over-parameterized matrix sensing
Kelly Geyer, Anastasios Kyrillidis, and Amir Kalev · 2020
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Later among the works it cites.
Kernel and Rich Regimes in Overparametrized Models
Blake Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2020
Later among the works it cites.
Robust recovery via implicit bias of discrepant learning rates for double over-parameterization
Chong You, Zhihui Zhu, Qing Qu, and Yi Ma · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hedy Attouch, Jérôme Bolte, and Benar Fux Svaiter · 2013
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
A weighted ℓ 1 \ell_{1} -minimization approach for sparse polynomial chaos expansions
Ji Peng, Jerrad Hampton, and Alireza Doostan · 2014
Cited alongside, same era.
Projected gradient method for non-negative least square
Roman A Polyak · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
Accelerated mirror descent in continuous and discrete time
Walid Krichene, Alexandre Bayen, and Peter L Bartlett · 2015
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
Dominik Stöger and Mahdi Soltanolkotabi · 2021
Later among the works it cites.
Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
On the implicit bias of initialization shape: Beyond infinitesimal mirror descent
Shahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E Woodworth, Nathan Srebro, Amir Globerson, and Daniel Soudry · 2021
Later among the works it cites.
Continuous vs. discrete optimization of deep neural networks
Omer Elkabetz and Nadav Cohen · 2021
Later among the works it cites.
Enhanced Resolution Analysis for Water Molecules in MCM-41 and SBA-15 in Low-Field T2 Relaxometric Spectra
Grzegorz Stoch and Artur T Krzyżak · 2021
Later among the works it cites.
An alternating rank-k nonnegative least squares framework (ARkNLS) for nonnegative matrix factorization
Delin Chu, Weya Shi, Srinivas Eswar, and Haesun Park · 2021
Later among the works it cites.
Implicit sparse regularization: The impact of depth and early stopping
Jiangyuan Li, Thanh Nguyen, Chinmay Hegde, and Ka Wai Wong · 2021
Later among the works it cites.
Implicit Regularization in Matrix Sensing via Mirror Descent
Fan Wu and Patrick Rebeschini · 2021
Later among the works it cites.
Mirrorless mirror descent: A natural derivation of mirror descent
Suriya Gunasekar, Blake Woodworth, and Nathan Srebro · 2021
Later among the works it cites.
On the best choice of LASSO program given data parameters
Aaron Berk, Yaniv Plan, and Ozgur Yilmaz · 2021
Later among the works it cites.
The lawson-hanson algorithm with deviation maximization: Finite convergence and sparse recovery
Monica Dessole, Marco Dell’Orto, and Fabio Marcuzzi · 2021
Later among the works it cites.
Implicit bias of gradient descent on reparametrized models: On equivalence to mirror descent
Zhiyuan Li, Tianhao Wang, Jason D Lee, and Sanjeev Arora · 2022
Closest in time.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Closest in time.
Large Learning Rate Tames Homogeneity: Convergence and Balancing Effect
Yuqing Wang, Minshuo Chen, Tuo Zhao, and Molei Tao · 2022
Closest in time.
Robust training under label noise by over-parameterization
Sheng Liu, Zhihui Zhu, Qing Qu, and Chong You · 2022
Closest in time.
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Closest in time.
From the simplex to the sphere: faster constrained optimization using the hadamard parametrization
Qiuwei Li, Daniel McKenzie, and Wotao Yin · 2023
Closest in time.
Accelerated riemannian optimization: Handling constraints with a prox to bound geometric penalties
David Martínez-Rubio and Sebastian Pokutta · 2023
Closest in time.
On squared-variable formulations
Lijun Ding and Stephen J Wright · 2023
Closest in time.
Smooth over-parameterized solvers for non-smooth structured optimization
Clarice Poon and Gabriel Peyré · 2023
Closest in time.
Incremental learning in diagonal linear networks
Raphaël Berthier · 2023
Closest in time.
Saddle-to-saddle dynamics in diagonal linear networks
Scott Pesme and Nicolas Flammarion · 2023
Closest in time.
Implicit bias of gradient descent for logistic regression at the edge of stability
Jingfeng Wu, Vladimir Braverman, and Jason D. Lee · 2023
Closest in time.
Yuqing Wang, Zhenghao Xu, Tuo Zhao, and Molei Tao · 2023
Closest in time.
An introduction to optimization on smooth manifolds
Nicolas Boumal · 2023
Closest in time.
Gradient descent for deep matrix factorization: Dynamics and implicit bias towards low rank
Hung-Hsu Chou, Carsten Gieshoff, Johannes Maly, and Holger Rauhut · 2024
Closest in time.
Co673/cs794 - optimization for data science lecture notes, 2017
Yao-Liang Yu · 2024
Closest in time.