Fetching the paper…
Reading the bibliography…
We introduce a technique for tuning the learning rate scale factor of any base optimization algorithm and schedule automatically, which we call \textsc{mechanic}.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Online Learning: Theory, Algorithms, and Applications
S. Shalev-Shwartz · 2007
Earlier work this paper cites.
Optimal strategies and minimax lower bounds for online convex games
Jacob Abernethy, Peter L Bartlett, Alexander Rakhlin, and Ambuj Tewari · 2008
Earlier work this paper cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2008
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2010
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. Brendan McMahan and Matthew Streeter · 2010
Earlier work this paper cites.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2011
Earlier work this paper cites.
No-regret algorithms for unconstrained online convex optimization
Brendan Mcmahan and Matthew Streeter · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Minimax optimal algorithms for unconstrained linear optimization
Brendan McMahan and Jacob Abernethy · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Muprop: Unbiased backpropagation for stochastic neural networks
Shixiang Shane Gu, Sergey Levine, Ilya Sutskever, and Andriy Mnih · 2015
Earlier work this paper cites.
Online convex optimization with unconstrained domains and losses
Ashok Cutkosky and Kwabena A Boahen · 2016
Earlier work this paper cites.
Coin betting and parameter-free online learning
Francesco Orabona and Dávid Pál · 2016
Earlier work this paper cites.
Online learning without prior information
Ashok Cutkosky and Kwabena Boahen · 2017
Cited alongside, same era.
Training deep networks without learning rates through coin betting
Francesco Orabona and Tatiana Tommasi · 2017
Cited alongside, same era.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Cited alongside, same era.
Automatic differentiation in machine learning: a survey
Atilim Gunes Baydin, Barak A Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind · 2018
Cited alongside, same era.
Black-box reductions for parameter-free online learning in Banach spaces
Ashok Cutkosky and Francesco Orabona · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
A modern introduction to online learning
Francesco Orabona · 2019
Later among the works it cites.
Disentangling adaptive gradient methods from learning rates
Naman Agarwal, Rohan Anil, Elad Hazan, Tomer Koren, and Cyril Zhang · 2020
Later among the works it cites.
Lipschitz and comparator-norm adaptivity in online learning
Zakaria Mhammedi and Wouter M Koolen · 2020
Later among the works it cites.
Impossible tuning made possible: A new expert algorithm and its applications
Liyu Chen, Haipeng Luo, and Chen-Yu Wei · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2018
Cited alongside, same era.
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Nathan Srebro, and Blake Woodworth · 2019
Cited alongside, same era.
Artificial constraints and hints for unbounded online learning
Ashok Cutkosky · 2019
Cited alongside, same era.
Combining online learning guarantees
Ashok Cutkosky · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Introduction to online convex optimization
Elad Hazan · 2019
Cited alongside, same era.
Storm+: Fully adaptive sgd with recursive momentum for nonconvex optimization
Kfir Levy, Ali Kavis, and Volkan Cevher · 2021
Later among the works it cites.
Making sgd parameter-free
Yair Carmon and Oliver Hinder · 2022
Later among the works it cites.
Gradient descent: The ultimate optimizer
Kartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, and Erik Meijer · 2022
Later among the works it cites.
Parameter-free mirror descent
Andrew Jacobsen and Ashok Cutkosky · 2022
Later among the works it cites.
Adaptive gradient methods with local guarantees
Zhou Lu, Wenhan Xia, Sanjeev Arora, and Elad Hazan · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, et al · 2023
Closest in time.
Symbolic discovery of optimization algorithms, 2023
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, and Quoc V. Le · 2023
Closest in time.
A nonstochastic control approach to optimization
Xinyi Chen and Elad Hazan · 2023
Closest in time.
Learning-rate-free learning by d-adaptation
Aaron Defazio and Konstantin Mishchenko · 2023
Closest in time.
Dog is sgd’s best friend: A parameter-free dynamic step size schedule
Maor Ivgi, Oliver Hinder, and Yair Carmon · 2023
Closest in time.