Fetching the paper…
Reading the bibliography…
Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T.
A modern introduction to online learning
Orabona, F. (2019) · 1912
Earlier work this paper cites.
A method for solving a convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Nesterov, Y. (1983) · 1983
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins-Monro process
Ruppert, D. (1988) · 1988
Earlier work this paper cites.
New stochastic approximation type procedures
Polyak, B. (1990) · 1990
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M. (2003) · 2003
Earlier work this paper cites.
Extracting certainty from uncertainty: Regret bounded by variation in costs
Hazan, E. and Kale, S. (2010) · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
Online optimization with gradual variations
Chiang, C.-K., Yang, T., Lee, C.-J., Mahdavi, M., Lu, C.-J., Jin, R., and Zhu, S. (2012) · 2012
Earlier work this paper cites.
A simpler approach to obtaining an o ( 1 / t ) o(1/t) convergence rate for the projected stochastic subgradient method
Lacoste-Julien, S., Schmidt, M., and Bach, F. (2012) · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Lan, G. (2012) · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K. (2012) · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) O(1/n)
Bach, F. and Moulines, E. (2013) · 2013
Earlier work this paper cites.
Lectures on Convex Optimization
Nesterov, Y. (2013) · 2013
Earlier work this paper cites.
Online learning with predictable sequences
Rakhlin, A. and Sridharan, K. (2013) · 2013
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Shamir, O. and Zhang, T. (2013) · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. E. (2013) · 2013
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Cettolo, M., Niehues, J., Stüker, S., Bentivogli, L., and Federico, M. (2014) · 2014
Earlier work this paper cites.
Display advertising challenge
Jean-Baptiste Tien, joycenv, O. C. (2014) · 2014
Earlier work this paper cites.
Adam: a method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Quasi-monotone subgradient methods for nonsmooth convex minimization
Nesterov, Y. and Shikhman, V. (2015) · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. (2015) · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. (2015) · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016) · 2016
Cited alongside, same era.
Sequence-to-sequence learning as beam-search optimization
Wiseman, S. and Rush, A. M. (2016) · 2016
Cited alongside, same era.
Wide residual networks
Zagoruyko, S. and Komodakis, N. (2016) · 2016
Cited alongside, same era.
Findings of the 2017 conference on machine translation (wmt17)
Bojar, O., Chatterjee, R., Federmann, C., Graham, Y., Haddow, B., Huang, S., Huck, M., Koehn, P., Liu, Q., Logacheva, V., Monz, C., Negri, M., Post, M., Rubino, R., Specia, L., and Turchi, M. (2017) · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q. (2017) · 2017
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition
Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., and Pang, R. (2020) · 2020
Later among the works it cites.
Open graph benchmark: datasets for machine learning on graphs
Hu, W., Fey, M., Zitnik, M., Dong, Y., Ren, H., Liu, B., Catasta, M., and Leskovec, J. (2020) · 2020
Later among the works it cites.
A simpler approach to accelerated optimization: iterative averaging meets optimism
Joulani, P., Raj, A., Gyorgy, A., and Szepesvári, C. (2020) · 2020
Later among the works it cites.
End-to-end variational networks for accelerated MRI reconstruction
Sriram, A., Zbontar, J., Murrell, T., Defazio, A., Zitnick, C. L., Yakubova, N., Knoll, F., and Johnson, P. (2020) · 2020
Later among the works it cites.
The power of factorial powers: New parameter settings for (stochastic) optimization
Defazio, A. and Gower, R. M. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A modular analysis of adaptive (non-) convex optimization: Optimism, composite objectives, and variational bounds
Joulani, P., György, A., and Szepesvári, C. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Relational inductive biases, deep learning, and graph networks
Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., Gulcehre, C., Song, F., Ballard, A., Gilmer, J., Dahl, G., Vaswani, A., Allen, K., Nash, C., Langston, V., Dyer, C., Heess, N., Wierstra, D., Kohli, P., Botvinick, M., Vinyals, O., Li, Y., and Pascanu, R. (2018) · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G. (2018) · 2018
Cited alongside, same era.
On the convergence of Adam and beyond
Reddi, S. J., Kale, S., and Kumar, S. (2018) · 2018
Cited alongside, same era.
Primal averaging: A new gradient evaluation step to attain the optimal individual convergence
Tao, W., Pan, Z., Wu, G., and Tao, Q. (2018) · 2018
Cited alongside, same era.
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. (2021) · 2021
Later among the works it cites.
On the (asymptotic) convergence of stochastic gradient descent and stochastic heavy ball
Sebbouh, O., Gower, R. M., and Defazio, A. (2021) · 2021
Later among the works it cites.
Rethinking "batch" in batchnorm
Wu, Y. and Johnson, J. (2021) · 2021
Later among the works it cites.
Criteo 1TB click logs dataset
Criteo (2022) · 2022
Later among the works it cites.
Adaptivity without compromise: A momentumized, adaptive, dual averaged gradient method for stochastic optimization
Defazio, A. and Jelassi, S. (2022) · 2022
Later among the works it cites.
Introduction to online convex optimization
Hazan, E. (2022) · 2022
Later among the works it cites.
Stop wasting my time! saving days of ImageNet and BERT training with latest weight averaging
Kaddour, J. (2022) · 2022
Later among the works it cites.
Fast benchmarking of accuracy vs. training time with cyclic learning rates
Portes, J., Blalock, D., Stephenson, C., and Frankle, J. (2022) · 2022
Later among the works it cites.
Benchmarking Neural Network Training Algorithms
Dahl, G. E., Schneider, F., Nado, Z., Agarwal, N., Sastry, C. S., Hennig, P., Medapati, S., Eschenhagen, R., Kasimbeg, P., Suo, D., Bae, J., Gilmer, J., Peirson, A. L., Khan, B., Anil, R., Rabbat, M., Krishnan, S., Snider, D., Amid, E., Chen, K., Maddison, C. J., Vasudev, R., Badura, M., Garg, A., and Mattson, P. (2023) · 2023
Later among the works it cites.
When, why and how much? adaptive learning rate scheduling by refinement
Defazio, A., Cutkosky, A., Mehta, H., and Mishchenko, K. (2023) · 2023
Later among the works it cites.
Learning-rate-free learning by D-adaptation
Defazio, A. and Mishchenko, K. (2023) · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., Jenatton, R., Beyer, L., Tschannen, M., Arnab, A., Wang, X., Riquelme Ruiz, C., Minderer, M., Puigcerver, J., Evci, U., Kumar, M., Steenkiste, S. V., Elsayed, G. F., Mahendran, A., Yu, F., Oliver, A., Huot, F., Bastings, J., Collier, M., Gritsenko, A. A., Birodkar, V., Vasconcelos, C. N., Tay, Y., Mensink, T., Kolesnikov, A., Pavetic, F., Tran, D., Kipf, T., Lucic, M., Zhai, X., Keysers, D., Harmsen, J. J., and Houlsby, N. (2023) · 2023
Later among the works it cites.
Training trajectories, mini-batch losses and the curious role of the learning rate
Sandler, M., Zhmoginov, A., Vladymyrov, M., and Miller, N. (2023) · 2023
Later among the works it cites.
Early weight averaging meets high learning rates for LLM pre-training
Sanyal, S., Neerkaje, A., Kaddour, J., Kumar, A., and Sanghavi, S. (2023) · 2023
Later among the works it cites.
Exact convergence rate of the last iterate in subgradient methods
Zamani, M. and Glineur, F. (2023) · 2023
Later among the works it cites.
On the generalization ability of on-line learning algorithms
Cesa-Bianchi, N., Conconi, A., and Gentile, C. (2004) · 2057
Closest in time.