Fetching the paper…
Reading the bibliography…
Learning rate schedules used in practice bear little resemblance to those recommended by theory.
RoBERTa: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
Learning an adaptive learning rate schedule
Xu, Z., Dai, A. M., Kemp, J., and Metz, L. (2019) · 1909
Earlier work this paper cites.
On the determination of the step size in stochastic quasigradient methods
Pflug, G. C. (1983) · 1983
Earlier work this paper cites.
Adaptive stepsize control in stochastic approximation algorithms
Pflug, G. C. (1988) · 1988
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
Ruppert, D. (1988) · 1988
Earlier work this paper cites.
New stochastic approximation type procedures
Polyak, B. (1990) · 1990
Earlier work this paper cites.
Parameter Adaptation in Stochastic Optimization
Almeida, L. B., Langlois, T., Amaral, J. D., and Plakhov, A. (1999) · 1999
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Bengio, Y. (2000) · 2000
Earlier work this paper cites.
Statistical adaptive stochastic gradient methods
Zhang, P., Lang, H., Liu, Q., and Xiao, L. (2020) · 2002
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M. (2003) · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Earlier work this paper cites.
Less regret via online conditioning
Streeter, M. and McMahan, H. B. (2010) · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
Generic methods for optimization-based modeling
Domke, J. (2012) · 2012
Earlier work this paper cites.
A simpler approach to obtaining an O ( 1 / t ) O(1/t) convergence rate for the projected stochastic subgradient method
Lacoste-Julien, S., Schmidt, M., and Bach, F. (2012) · 2012
Earlier work this paper cites.
No-regret algorithms for unconstrained online convex optimization
Mcmahan, B. and Streeter, M. (2012) · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K. (2012) · 2012
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Shamir, O. and Zhang, T. (2013) · 2013
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Cettolo, M., Niehues, J., Stüker, S., Bentivogli, L., and Federico, M. (2014) · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J. (2015) · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. (2015) · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. (2015) · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Iterative regularization for learning with convex loss functions
Lin, J., Rosasco, L., and Zhou, D.-X. (2016) · 2016
Cited alongside, same era.
Coin betting and parameter-free online learning
Orabona, F. and Pál, D. (2016) · 2016
Cited alongside, same era.
Hyperparameter optimization with approximate gradient
Pedregosa, F. (2016) · 2016
Cited alongside, same era.
Sequence-to-sequence learning as beam-search optimization
Wiseman, S. and Rush, A. M. (2016) · 2016
Cited alongside, same era.
Adagrad stepsizes: sharp convergence over nonconvex landscapes
Ward, R., Wu, X., and Bottou, L. (2019) · 2019
Later among the works it cites.
Marthe: Scheduling the learning rate via online hypergradients
Donini, M., Franceschi, L., Majumder, O., Pontil, M., and Frasconi, P. (2020) · 2020
Later among the works it cites.
The complexity of finding stationary points with stochastic gradient descent
Drori, Y. and Shamir, O. (2020) · 2020
Later among the works it cites.
Lipschitz and comparator-norm adaptivity in online learning
Mhammedi, Z. and Koolen, W. M. (2020) · 2020
Later among the works it cites.
Last iterate of SGD converges (even in unbounded domains)
Orabona, F. (2020) · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wide residual networks
Zagoruyko, S. and Komodakis, N. (2016) · 2016
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Franceschi, L., Donini, M., Frasconi, P., and Pontil, M. (2017) · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Goya, P., Dollar, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. (2017) · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q. (2017) · 2017
Cited alongside, same era.
Online learning rate adaptation with hypergradient descent
Baydin, A. G., Cornish, R., Rubio, D. M., Schmidt, M., and Wood, F. (2018) · 2018
Cited alongside, same era.
Black-box reductions for parameter-free online learning in banach spaces
Cutkosky, A. and Orabona, F. (2018) · 2018
Cited alongside, same era.
Later among the works it cites.
End-to-end variational networks for accelerated MRI reconstruction
Sriram, A., Zbontar, J., Murrell, T., Defazio, A., Zitnick, C. L., Yakubova, N., Knoll, F., and Johnson, P. (2020) · 2020
Later among the works it cites.
WNGrad: Learn the learning rate in gradient descent
Wu, X., Ward, R., and Bottou, L. (2020) · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. (2021) · 2021
Later among the works it cites.
Autolrs: Automatic learning-rate schedule by bayesian optimization on the fly
Jin, Y., Zhou, T., Zhao, L., Zhu, Y., Guo, C., Canini, M., and Krishnamurthy, A. (2021) · 2021
Later among the works it cites.
Automated learning rate scheduler for large-batch training
Kim, C., Kim, S., Kim, J., Lee, D., and Kim, S. (2021) · 2021
Later among the works it cites.
Making SGD parameter-free
Carmon, Y. and Hinder, O. (2022) · 2022
Later among the works it cites.
Gradient descent: The ultimate optimizer
Chandra, K., Xie, A., Ragan-Kelley, J., and Meijer, E. (2022) · 2022
Later among the works it cites.
An empirical analysis of compute-optimal large language model training
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J. W., and Sifre, L. (2022) · 2022
Later among the works it cites.
Eigencurve: Optimal learning rate schedule for SGD on quadratic objectives with skewed hessian spectrums
Pan, R., Ye, H., and Zhang, T. (2022) · 2022
Later among the works it cites.
PDE-based optimal strategy for unconstrained online learning
Zhang, Z., Cutkosky, A., and Paschalidis, I. (2022) · 2022
Later among the works it cites.
Mechanic: A learning rate tuner
Cutkosky, A., Defazio, A., and Mehta, H. (2023) · 2023
Closest in time.
DoG is SGD’s best friend: A parameter-free dynamic step size schedule
Ivgi, M., Hinder, O., and Carmon, Y. (2023) · 2023
Closest in time.
DoWG unleashed: An efficient universal parameter-free gradient descent method
Khaled, A., Mishchenko, K., and Jin, C. (2023) · 2023
Closest in time.
Exact convergence rate of the last iterate in subgradient methods
Zamani, M. and Glineur, F. (2023) · 2023
Closest in time.
On the generalization ability of on-line learning algorithms
Cesa-Bianchi, N., Conconi, A., and Gentile, C. (2004) · 2057
Closest in time.