Fetching the paper…
Reading the bibliography…
The backtracking line-search is an effective technique to automatically tune the step-size in smooth optimization.
“First-Order Preconditioning via Hypergradient Descent” arXiv/1910.08461
Ted Moskovitz, Rui Wang, Janice Lan, Sanyam Kapoor, Thomas Miconi, Jason Yosinski and Aditya Rawal · 1910
Earlier work this paper cites.
“Accelerated Stochastic Approximation”
Harry Kesten · 1958
Earlier work this paper cites.
“Une propriété topologique des sous-ensembles analytiques réels”
S. Łojasiewicz · 1963
Earlier work this paper cites.
“Gradient methods for minimizing functionals”
Boris. Polyak · 1963
Earlier work this paper cites.
“Minimization of functions having Lipschitz continuous first partial derivatives”
Larry Armijo · 1966
Earlier work this paper cites.
“Convergence conditions for ascent methods”
Philip Wolfe · 1969
Earlier work this paper cites.
“Learning Applied to Successive Approximation Algorithms”
George. Saridis · 1970
Earlier work this paper cites.
“Informational complexity and effective methods for the solution of convex extremal problems”
David. Yudin and Arkadi. Nemirovski · 1976
Earlier work this paper cites.
“Quasi-Newton methods, motivation and theory”
John. Dennis Jr. and Jorge. Moré · 1977
Earlier work this paper cites.
“Cut-off method with space extension in convex programming problems”
Naum. Shor · 1977
Earlier work this paper cites.
“Goal Seeking Components for Adaptive Intelligence: An Initial Assessment.” (Appendix C), 1981
Andrew. Barto and Richard. Sutton · 1981
Earlier work this paper cites.
“A Nonmonotone Line Search Technique for Newton’s Method”
Luigi Grippo, Francesco Lampariello and Stephano Lucidi · 1986
Earlier work this paper cites.
“Increased rates of convergence through learning rate adaptation”
Robert. Jacobs · 1988
Earlier work this paper cites.
“On the limited memory BFGS method for large scale optimization”
Dong. Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
“Acceleration techniques for the backpropagation algorithm”
Fernando. Silva and Luís. Almeida · 1990
Earlier work this paper cites.
“Adapting Bias by Gradient Descent: An Incremental Version of Delta-Bar-Delta”
Richard. Sutton · 1992
Earlier work this paper cites.
“Gain adaptation beats least squares”
Richard. Sutton · 1992
Earlier work this paper cites.
“A direct adaptive method for faster backpropagation learning: the RPROP algorithm”
Martin. Riedmiller and Heinrich Braun · 1993
Earlier work this paper cites.
“Line search algorithms with guaranteed sufficient decrease”
Jorge. Moré and David. Thuente · 1994
Earlier work this paper cites.
“Sparse spatial autoregressions”
R. Kelley Pace and Ronald Barry · 1997
Earlier work this paper cites.
“Algorithm 778: L-BFGS-B: Fortran Subroutines for Large-Scale Bound-Constrained Optimization”
Ciyou Zhu, Richard. Byrd, Peihuang Lu and Jorge Nocedal · 1997
Earlier work this paper cites.
“Modeling of strength of high-performance concrete using artificial neural networks”
I.-Cheng Yeh · 1998
Earlier work this paper cites.
“Parameter adaptation in stochastic optimization”
Luís. Almeida, Thibault Langlois, José.. Amaral and Alexander Plakhov · 1999
Earlier work this paper cites.
“Nonlinear Programming”
Dimitri. Bertsekas · 1999
Earlier work this paper cites.
“Numerical Optimization”
Jorge Nocedal and Stephen. Wright · 1999
Earlier work this paper cites.
“Local gain adaptation in stochastic gradient descent”
Nicol. Schraudolph · 1999
Earlier work this paper cites.
“The Quasi-Cauchy Relation and Diagonal Updating”
M. Zhu, John. Nazareth and Henry Wolkowicz · 1999
Cited alongside, same era.
“Learning rate adaptation in stochastic gradient descent”
Vassilis. Plagianakos, George. Magoulas and Michael. Vrahatis · 2001
Cited alongside, same era.
“Disentangling Adaptive Gradient Methods from Learning Rates” arXiv/2002.11803
Naman Agarwal, Rohan Anil, Elad Hazan, Tomer Koren and Cyril Zhang · 2002
Cited alongside, same era.
“Efficient SVM Regression Training with SMO”
Gary Flake and Steve Lawrence · 2002
Cited alongside, same era.
“RCV1: A New Benchmark Collection for Text Categorization Research”
David. Lewis, Yiming Yang, Tony. Rose and Fan Li · 2004
Cited alongside, same era.
“UCI Machine Learning Repository”, 2017
Dheeru Dua and Casey Graff · 2017
Later among the works it cites.
“On the worst-case complexity of the gradient method with exact line search for smooth strongly convex functions”
Etienne de Klerk, François Glineur and Adrien. Taylor · 2017
Later among the works it cites.
“Training Deep Networks without Learning Rates Through Coin Betting”
Francesco Orabona and Tatiana Tommasi · 2017
Later among the works it cites.
“Online Learning Rate Adaptation with Hypergradient Descent”
Atılımüneş Baydin, Robert Cornish, David Martínez-Rubio, Mark Schmidt and Frank Wood · 2018
Later among the works it cites.
“A Progressive Batching L-BFGS Method for Machine Learning”
Raghu Bollapragada, Dheevatsa Mudigere, Jorge Nocedal, Hao-Jun Shi and Ping Tang · 2018
Later among the works it cites.
“A diagonal quasi-Newton updating method for unconstrained optimization”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“A Modified Finite Newton Method for Fast Solution of Large Scale Linear SVMs”
S. Keerthi and Dennis DeCoste · 2005
Cited alongside, same era.
“Fast Online Policy Gradient Learning with SMD Gain Vector Adaptation”
Nicol. Schraudolph, Douglas Aberdeen and Jin Yu · 2005
Cited alongside, same era.
“Cubic regularization of Newton method and its global performance”
Yurii. Nesterov and Boris. Polyak · 2006
Cited alongside, same era.
“Approximating submodular functions everywhere”
Michel. Goemans, Nicholas.. Harvey, Satoru Iwata and Vahab Mirrokni · 2009
Cited alongside, same era.
“Adaptive Bound Optimization for Online Convex Optimization”
H. McMahan and Matthew. Streeter · 2010
Cited alongside, same era.
“LIBSVM: A Library for Support Vector Machines”
Chih-Chung Chang and Chih-Jen Lin · 2011
Cited alongside, same era.
“Adaptive Subgradient Methods for Online Learning and Stochastic Optimization”
John. Duchi, Elad Hazan and Yoram Singer · 2011
Cited alongside, same era.
Neculai Andrei · 2019
Later among the works it cites.
“On the Convergence of Stochastic Gradient Descent with Adaptive Stepsizes”
Xiaoyu Li and Francesco Orabona · 2019
Later among the works it cites.
“PyTorch: An Imperative Style, High-Performance Deep Learning Library”
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai and Soumith Chintala · 2019
Later among the works it cites.
“Painless Stochastic Gradient: Interpolation, Line-Search, and Convergence Rates”
Sharan Vaswani, Aaron Mishkin, Issam. Laradji, Mark Schmidt, Gauthier Gidel and Simon Lacoste-Julien · 2019
Later among the works it cites.
“AdaGrad stepsizes: sharp convergence over nonconvex landscapes”
Rachel Ward, Xiaoxia Wu and Léon Bottou · 2019
Later among the works it cites.
“Fast and Near-Optimal Diagonal Preconditioning” arXiv/2008.01722, 2020
Arun Jambulapati, Jerry Li, Christopher Musco, Aaron Sidford and Kevin Tian · 2020
Later among the works it cites.
“Fast and Furious Convergence: Stochastic Second Order Methods under Interpolation”
Si Meng, Sharan Vaswani, Issam Laradji, Mark Schmidt and Simon Lacoste-Julien · 2020
Later among the works it cites.
“Variable Metric Proximal Gradient Method with Diagonal Barzilai-Borwein Stepsize”
Youngsuk Park, Sauptik Dhar, Stephen. Boyd and Mohak Shah · 2020
Later among the works it cites.
“Adaptive Gradient Methods Converge Faster with Over-Parameterization (and you can do a line-search)” arXiv/2006.06835, 2020
Sharan Vaswani, Frederik Kunstner, Issam. Laradji, Si Meng, Mark Schmidt and Simon Lacoste-Julien · 2020
Later among the works it cites.
“SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python”
Pauli Virtanen, Ralf Gommers, Travis. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan. van der Walt, Matthew Brett, Joshua Wilson, K. Millman, Nikolay Mayorov, Andrew.. Nelson, Eric Jones, Robert Kern, Eric Larson, C Carey, İlhan Polat, Yu Feng, Eric. Moore, Jake VanderPlas, Denis Laxalde, Josef Perktold, Robert Cimrman, Ian Henriksen, E.. Quintero, Charles. Harris, Anne. Archibald, Antônio. Ribeiro, Fabian Pedregosa, Paul van Mulbregt and SciPy 1.0 Contributors · 2020
Later among the works it cites.
“Quasi-Newton methods for machine learning: forget the past, just sample”
Albert. Berahas, Majid Jahani, Peter Richtárik and Martin Takáč · 2021
Later among the works it cites.
“Smoothness Matrices Beat Smoothness Constants: Better Communication Compression Techniques for Distributed Optimization”
Mher Safaryan, Filip Hanzely and Peter Richtárik · 2021
Later among the works it cites.
“AdaHessian: An Adaptive Second Order Optimizer for Machine Learning”
Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer and Michael Mahoney · 2021
Later among the works it cites.
Ehsan Amid, Rohan Anil, Christopher Fifty and Manfred. Warmuth · 2022
Later among the works it cites.
“Amortized Proximal Optimization”
Juhan Bae, Paul Vicol, Jeff. HaoChen and Roger. Grosse · 2022
Later among the works it cites.
“Gradient Descent: The Ultimate Optimizer”
Kartik Chandra, Audrey Xie, Jonathan Ragan-Kelley and Erik Meijer · 2022
Later among the works it cites.
“Grad-GradaGrad? A Non-Monotone Adaptive Stochastic Gradient Method” arXiv/2206.06900
Aaron Defazio, Baoyu Zhou and Lin Xiao · 2022
Later among the works it cites.
“A Simple Convergence Proof of Adam and Adagrad”
Alexandre Défossez, Leon Bottou, Francis Bach and Nicolas Usunier · 2022
Later among the works it cites.
“Doubly Adaptive Scaled Algorithm for Machine Learning Using Second-Order Information”
Majid Jahani, Sergey Rusakov, Zheng Shi, Peter Richtárik, Michael. Mahoney and Martin Takac · 2022
Later among the works it cites.
“Optimal Diagonal Preconditioning: Theory and Practice” arXiv/2209.00809, 2022
Zhaonan Qu, Wenzhi Gao, Oliver Hinder, Yinyu Ye and Zhengyuan Zhou · 2022
Later among the works it cites.