Fetching the paper…
Reading the bibliography…
We consider the problem of estimating the learning rate in adaptive methods, such as AdaGrad and Adam.
Introduction to optimization
Boris T. Polyak · 1987
Earlier work this paper cites.
Support vector machines for multi-class pattern recognition
Jason Weston and Christopher Watkins · 1999
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Less regret via online conditioning
Matthew Streeter and H. Brendan McMahan · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Unconstrained online linear learning in Hilbert spaces: Minimax algorithms and normal approximations
H. Brendan McMahan and Francesco Orabona · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Training deep networks without learning rates through coin betting
Francesco Orabona and Tatiana Tommasi · 2017
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc Le · 2017
Earlier work this paper cites.
Black-box reductions for parameter-free online learning in banach spaces
Ashok Cutkosky and Francesco Orabona · 2018
Earlier work this paper cites.
Online adaptive methods, universality and acceleration
Kfir Y. Levy, Alp Yurtsever, and Volkan Cevher · 2018
Cited alongside, same era.
fastMRI: An open dataset and benchmarks for accelerated MRI
Jure Zbontar, Florian Knoll, Anuroop Sriram, Matthew J. Muckley, Mary Bruno, Aaron Defazio, Marc Parente, Krzysztof J. Geras, Joe Katsnelson, Hersh Chandarana, et al · 2018
Cited alongside, same era.
Revisiting the Polyak step size
Elad Hazan and Sham M. Kakade · 2019
Cited alongside, same era.
UniXGrad: A universal, adaptive algorithm with optimal guarantees for constrained optimization
Ali Kavis, Kfir Y. Levy, Francis Bach, and Volkan Cevher · 2019
Cited alongside, same era.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Advances and open problems in federated learning
Peter Kairouz, H. Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al · 2021
Later among the works it cites.
Federated hyperparameter tuning: Challenges, baselines, and connections to weight-sharing
Mikhail Khodak, Renbo Tu, Tian Li, Liam Li, Maria-Florina F. Balcan, Virginia Smith, and Ameet Talwalkar · 2021
Later among the works it cites.
Stochastic Polyak step-size for SGD: An adaptive learning rate for fast convergence
Nicolas Loizou, Sharan Vaswani, Issam Laradji, and Simon Lacoste-Julien · 2021
Later among the works it cites.
Parameter-free stochastic optimization of variationally coherent functions, 2021
Francesco Orabona and Dávid Pál · 2021
Later among the works it cites.
Making SGD parameter-free
Yair Carmon and Oliver Hinder · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep learning recommendation model for personalization and recommendation systems
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherniavskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, Volodymyr Kondratenko, Stephanie Pereira, Xianjie Chen, Wenlin Chen, Vijay Rao, Bill Jia, Liang Xiong, and Misha Smelyanskiy · 2019
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2019
Cited alongside, same era.
Adagrad stepsizes: sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Cited alongside, same era.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Cited alongside, same era.
Adaptive gradient descent without descent
Yura Malitsky and Konstantin Mishchenko · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Stochastic Polyak stepsize with a moving target
Robert M. Gower, Aaron Defazio, and Michael Rabbat · 2021
Cited alongside, same era.
Aaron Defazio and Samy Jelassi · 2022
Later among the works it cites.
Dynamics of SGD with stochastic Polyak stepsizes: Truly adaptive variants and convergence to exact solution
Antonio Orvieto, Simon Lacoste-Julien, and Nicolas Loizou · 2022
Later among the works it cites.
PDE-based optimal strategy for unconstrained online learning
Zhiyu Zhang, Ashok Cutkosky, and Ioannis Ch. Paschalidis · 2022
Later among the works it cites.
Learning-rate-free learning by D-adaptation
Aaron Defazio and Konstantin Mishchenko · 2023
Closest in time.
DoG is SGD’s best friend: A parameter-free dynamic step size schedule
Maor Ivgi, Oliver Hinder, and Yair Carmon · 2023
Closest in time.
Puya Latafat, Andreas Themelis, Lorenzo Stella, and Panagiotis Patrinos · 2023
Closest in time.