Fetching the paper…
Reading the bibliography…
D-Adaptation is an approach to automatically setting the learning rate which asymptotically achieves the optimal rate of convergence for minimizing convex Lipschitz functions, with no back-tracking or line searches, and no additional function value or gradient evaluations per step.
Introduction to optimization
Boris T. Polyak · 1987
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Yoshua Bengio · 2000
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Less regret via online conditioning, 2010
Matthew Streeter and H. Brendan McMahan · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Generic methods for optimization-based modeling
Justin Domke · 2012
Earlier work this paper cites.
No-regret algorithms for unconstrained online convex optimization
Matthew Streeter and H. Brendan McMahan · 2012
Earlier work this paper cites.
Introduction to Nonlinear Optimization
Amir Beck · 2014
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico · 2014
Earlier work this paper cites.
Unconstrained online linear learning in hilbert spaces: Minimax algorithms and normal approximations
H. Brendan McMahan and Francesco Orabona · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Hyperparameter optimization with approximate gradient
Fabian Pedregosa · 2016
Earlier work this paper cites.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M. Rush · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
Training deep networks without learning rates through coin betting
Francesco Orabona and Tatiana Tommasi · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Cited alongside, same era.
Lectures on Convex Optimization
Yurii Nesterov · 2018
Cited alongside, same era.
fastMRI: An open dataset and benchmarks for accelerated MRI
Jure Zbontar, Florian Knoll, Anuroop Sriram, Matthew J. Muckley, Mary Bruno, Aaron Defazio, Marc Parente, Krzysztof J. Geras, Joe Katsnelson, Hersh Chandarana, et al · 2018
Pytorch image models
Ross Wightman · 2019
Later among the works it cites.
Detectron2
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick · 2019
Later among the works it cites.
Momentum via primal averaging: Theoretical insights and learning rate schedules for non-convex optimization, 2020
Aaron Defazio · 2020
Later among the works it cites.
Efficient first-order methods for convex minimization: a constructive approach
Yoel Drori and Adrien B. Taylor · 2020
Later among the works it cites.
fastMRI: A publicly available raw k-space and DICOM dataset of knee images for accelerated MR image reconstruction using machine learning
Florian Knoll, Jure Zbontar, Anuroop Sriram, Matthew J. Muckley, Mary Bruno, Aaron Defazio, Marc Parente, Krzysztof J. Geras, Joe Katsnelson, Hersh Chandarana, Zizhao Zhang, Michal Drozdzalv, Adriana Romero, Michael Rabbat, Pascal Vincent, James Pinkerton, Duo Wang, Nafissa Yakubova, Erich Owens, C. Lawrence Zitnick, Michael P. Recht, Daniel K. Sodickson, and Yvonne W. Lui · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Offset sampling improves deep learning based accelerated mri reconstructions by exploiting symmetry
Aaron Defazio · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Automated Machine Learning , chapter Hyperparameter Optimization
Matthias Feurer and Frank Hutter · 2019
Cited alongside, same era.
Revisiting the polyak step size
Elad Hazan and Sham M. Kakade · 2019
Cited alongside, same era.
Making the last iterate of sgd information theoretically optimal
Prateek Jain, Dheeraj Nagaraj, and Praneeth Netrapalli · 2019
Cited alongside, same era.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
End-to-end variational networks for accelerated MRI reconstruction
Anuroop Sriram, Jure Zbontar, Tullie Murrell, Aaron Defazio, C. Lawrence Zitnick, Nafissa Yakubova, Florian Knoll, and Patricia Johnson · 2020
Later among the works it cites.
The power of factorial powers: New parameter settings for (stochastic) optimization
Aaron Defazio and Robert M. Gower · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Parameter-free stochastic optimization of variationally coherent functions, 2021
Francesco Orabona and Dávid Pál · 2021
Later among the works it cites.
Guarantees for tuning the step size using a learning-to-learn approach
Xiang Wang, Shuai Yuan, Chenwei Wu, and Rong Ge · 2021
Later among the works it cites.
Making SGD parameter-free
Yair Carmon and Oliver Hinder · 2022
Later among the works it cites.
Gradient descent: The ultimate optimizer
Kartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, and Erik Meijer · 2022
Later among the works it cites.
Optimal first-order methods for convex functions with a quadratic upper bound
Baptiste Goujaud, Adrien Taylor, and Aymeric Dieuleveut · 2022
Later among the works it cites.
Pde-based optimal strategy for unconstrained online learning
Zhiyu Zhang, Ashok Cutkosky, and Ioannis Ch. Paschalidis · 2022
Later among the works it cites.
DoG is SGD’s best friend: A parameter-free dynamic step size schedule, 2023
Maor Ivgi, Oliver Hinder, and Yair Carmon · 2023
Closest in time.