Fetching the paper…
Reading the bibliography…
Towards designing learned optimization algorithms that are usable beyond their training setting, we identify key principles that classical algorithms obey, but have up to now, not been used for Learning to Optimize (L2O).
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
The convergence of a class of double-rank minimization algorithms
Charles George Broyden · 1970
Earlier work this paper cites.
A new approach to variable metric algorithms
Roger Fletcher · 1970
Earlier work this paper cites.
A family of variable-metric methods derived by variational means
Donald Goldfarb · 1970
Earlier work this paper cites.
Variations on variable-metric methods
John Greenstadt · 1970
Earlier work this paper cites.
Conditioning of quasi-Newton methods for function minimization
David F Shanno · 1970
Earlier work this paper cites.
Quasi-Newton methods, motivation and theory
John E Dennis, Jr and Jorge J Moré · 1977
Earlier work this paper cites.
A sparse quasi-Newton update derived variationally with a nondiagonally weighted Frobenius norm
Philippe L Toint · 1981
Earlier work this paper cites.
Variable metric methods for constrained optimization
Michael JD Powell · 1983
Earlier work this paper cites.
The convergence of variable metric matrices in unconstrained optimization
Ge Ren-Pu and Michael JD Powell · 1983
Earlier work this paper cites.
Introduction to optimization
Boris T Polyak · 1987
Earlier work this paper cites.
Two-point step size gradient methods
Jonathan Barzilai and Jonathan M Borwein · 1988
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Convergence of quasi-Newton matrices generated by the symmetric rank one update
Andrew R Conn, Nicholas IM Gould, and Philippe L Toint · 1991
Earlier work this paper cites.
Heavy-ball method in nonconvex optimization problems
SK Zavriev and FV Kostyuk · 1993
Earlier work this paper cites.
Python reference manual
Guido Rossum · 1995
Earlier work this paper cites.
The lack of a priori distinctions between learning algorithms
David H Wolpert · 1996
Earlier work this paper cites.
Nonlinear programming
Dimitri P Bertsekas · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Newton-like methods
Roger Fletcher · 2000
Earlier work this paper cites.
A second-order gradient-like dissipative dynamical system with Hessian-driven damping.: Application to optimization and mechanics
Felipe Alvarez, Hedy Attouch, Jérôme Bolte, and Patrick Redont · 2002
Earlier work this paper cites.
Matplotlib: A 2D graphics environment
John D Hunter · 2007
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction , volume 2
Trevor Hastie, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman · 2009
Earlier work this paper cites.
Learning fast approximations of sparse coding
Karol Gregor and Yann LeCun · 2010
Earlier work this paper cites.
The NumPy array: a structure for efficient numerical computation
Stéfan van der Walt, Chris Colbert, and Gael Varoquaux · 2011
Earlier work this paper cites.
A quasi-Newton proximal splitting method
Stephen Becker and Jalal Fadili · 2012
Earlier work this paper cites.
Quasi-Newton methods: A new direction
Philipp Hennig and Martin Kiefel · 2013
Earlier work this paper cites.
Plug-and-play priors for model based reconstruction
Singanallur V Venkatakrishnan, Charles A Bouman, and Brendt Wohlberg · 2013
Cited alongside, same era.
Global convergence of the heavy-ball method for convex optimization
Euhanna Ghadimi, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, H. Brendan Shillingford, and Nando De Freitas · 2016
Greedy quasi-Newton methods with explicit superlinear convergence
Anton Rodomanov and Yurii Nesterov · 2021
Later among the works it cites.
Guarantees for tuning the step size using a learning-to-learn approach
Xiang Wang, Shuai Yuan, Chenwei Wu, and Rong Ge · 2021
Later among the works it cites.
First-order optimization algorithms via inertial systems with Hessian driven damping
Hedy Attouch, Zaki Chbani, Jalal Fadili, and Hassan Riahi · 2022
Later among the works it cites.
Learning to optimize: A primer and a benchmark
Tianlong Chen, Xiaohan Chen, Wuyang Chen, Howard Heaton, Jialin Liu, Zhangyang Wang, and Wotao Yin · 2022
Later among the works it cites.
Unrolled variational bayesian algorithm for image blind deconvolution
Yunshi Huang, Emilie Chouzenoux, and Jean-Christophe Pesquet · 2022
Later among the works it cites.
Towards understanding how momentum improves generalization in deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast convex optimization via inertial dynamics with Hessian driven damping
Hedy Attouch, Juan Peypouquet, and Patrick Redont · 2016
Cited alongside, same era.
Learning to optimize
Ke Li and Jitendra Malik · 2016
Cited alongside, same era.
Learning gradient descent: Better generalization and longer horizons
Kaifeng Lv, Shunhua Jiang, and Jian Li · 2017
Cited alongside, same era.
Learning proximal operators: Using denoising networks for regularizing inverse imaging problems
Tim Meinhardt, Michael Moller, Caner Hazirbas, and Daniel Cremers · 2017
Cited alongside, same era.
Information-geometric optimization algorithms: A unifying picture via invariance principles
Yann Ollivier, Ludovic Arnold, Anne Auger, and Nikolaus Hansen · 2017
Cited alongside, same era.
Learned optimizers that scale and generalize
Olga Wichrowska, Niru Maheswaranathan, Matthew W Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Nando Freitas, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Samy Jelassi and Yuanzhi Li · 2022
Later among the works it cites.
Sharpened quasi-Newton methods: Faster superlinear rate and larger local convergence neighborhood
Qiujiang Jin, Alec Koppel, Ketan Rajawat, and Aryan Mokhtari · 2022
Later among the works it cites.
Practical tradeoffs between memory, compute, and performance in learned optimizers
Luke Metz, C Daniel Freeman, James Harrison, Niru Maheswaranathan, and Jascha Sohl-Dickstein · 2022
Later among the works it cites.
Shida Wang, Jalal Fadili, and Peter Ochs · 2022
Later among the works it cites.
Tutorial on amortized optimization
Brandon Amos · 2023
Later among the works it cites.
Imaging with equivariant deep learning: From unrolled network design to fully unsupervised learning
Dongdong Chen, Mike Davies, Matthias J Ehrhardt, Carola-Bibiane Schönlieb, Ferdia Sherry, and Julián Tachella · 2023
Later among the works it cites.
Handbook of convergence theorems for (stochastic) gradient methods
Guillaume Garrigos and Robert M Gower · 2023
Later among the works it cites.
Transformer-based learned optimization
Erik Gärtner, Luke Metz, Mykhaylo Andriluka, C Daniel Freeman, and Cristian Sminchisescu · 2023
Later among the works it cites.
Safeguarded learned convex optimization
Howard Heaton, Xiaohan Chen, Zhangyang Wang, and Wotao Yin · 2023
Later among the works it cites.
Learning to combine quasi-Newton methods
Maojia Li, Jialin Liu, and Wotao Yin · 2023
Later among the works it cites.
Learning to optimize quasi-Newton methods
Isaac Liao, Rumen Dangovski, Jakob Nicolaus Foerster, and Marin Soljacic · 2023
Later among the works it cites.
Towards constituting mathematical structures for learning to optimize
Jialin Liu, Xiaohan Chen, Zhangyang Wang, Wotao Yin, and HanQin Cai · 2023
Later among the works it cites.
PAC-Bayesian learning of optimization algorithms
Michael Sucker and Peter Ochs · 2023
Later among the works it cites.
Continuous Newton-like methods featuring inertia and variable mass
Camille Castera, Hedy Attouch, Jalal Fadili, and Peter Ochs · 2024
Closest in time.
What functions can graph neural networks compute on random graphs? the role of positional encoding
Nicolas Keriven and Samuel Vaiter · 2024
Closest in time.
Any-dimensional equivariant neural networks
Eitan Levin and Mateo Díaz · 2024
Closest in time.
Learning to optimize with convergence guarantees using nonlinear system theory
Andrea Martin and Luca Furieri · 2024
Closest in time.
Learning to warm-start fixed-point optimization algorithms
Rajiv Sambharya, Georgina Hall, Brandon Amos, and Bartolomeo Stellato · 2024
Closest in time.
Boosting data-driven mirror descent with randomization, equivariance, and acceleration
Hong Ye Tan, Subhadip Mukherjee, Junqi Tang, and Carola-Bibiane Schönlieb · 2024
Closest in time.
Equivariant plug-and-play image reconstruction
Matthieu Terris, Thomas Moreau, Nelly Pustelnik, and Julian Tachella · 2024
Closest in time.
Optimization dynamics of equivariant and augmented neural networks
Oskar Nordenfors, Fredrik Ohlsson, and Axel Flinth · 2025
Closest in time.