Fetching the paper…
Reading the bibliography…
We show that learning-rate schedules for large model training behave surprisingly similar to a performance bound from non-smooth convex optimization theory.
Convex Analysis
Rockafellar, R. T · 1970
Earlier work this paper cites.
Optimization and nonsmooth analysis
Clarke, F. H · 1983
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A. and Teboulle, M · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Zhang, T · 2004
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Shamir, O. and Zhang, T · 2013
Earlier work this paper cites.
Performance of first-order methods for smooth convex minimization: a novel approach
Drori, Y. and Teboulle, M · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
SGDR: stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Gradient descent learns linear dynamical systems
Hardt, M., Ma, T., and Recht, B · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Pytorch image models
Wightman, R · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning
Bolte, J. and Pauwels, E · 2021
Cited alongside, same era.
Making the last iterate of sgd information theoretically optimal
Jain, P., Nagaraj, D. M., and Netrapalli, P · 2021
Cited alongside, same era.
A second look at exponential and cosine step sizes: Simplicity, adaptivity, and performance
Li, X., Zhuang, Z., and Orabona, F · 2021
Cited alongside, same era.
Exact convergence rate of the last iterate in subgradient methods
Zamani, M. and Glineur, F · 2023
Later among the works it cites.
Why you don’t overfit, and don’t need Bayes if you only train for one epoch
Aitchison, L · 2024
Later among the works it cites.
Chinchilla scaling: A replication attempt
Besiroglu, T., Erdil, E., Barnett, M., and You, J · 2024
Later among the works it cites.
PEPit: computer-assisted worst-case analyses of first-order optimization methods in Python
Goujaud, B., Moucer, C., Glineur, F., Hendrickx, J. M., Taylor, A. B., and Dieuleveut, A · 2024
Later among the works it cites.
Scaling laws and compute-optimal training beyond fixed training durations
Hägele, A., Bakouch, E., Kosson, A., Ben allal, L., Von Werra, L., and Jaggi, M · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J., and Sifre, L · 2022
Cited alongside, same era.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2022
Cited alongside, same era.
Optimal linear decay learning rate schedules and further refinements
Defazio, A., Cutkosky, A., Mehta, H., and Mishchenko, K · 2023
Cited alongside, same era.
(S)GD over diagonal linear networks: Implicit bias, large stepsizes and edge of stability
Even, M., Pesme, S., Gunasekar, S., and Flammarion, N · 2023
Cited alongside, same era.
Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be
Kunstner, F., Chen, J., Lavington, J. W., and Schmidt, M · 2023
Cited alongside, same era.
Aiming towards the minimizers: fast convergence of sgd for overparametrized problems
Liu, C., Drusvyatskiy, D., Belkin, M., Davis, D., and Ma, Y · 2023
Cited alongside, same era.
Training trajectories, mini-batch losses and the curious role of the learning rate
Sandler, M., Zhmoginov, A., Vladymyrov, M., and Miller, N · 2023
Cited alongside, same era.
Later among the works it cites.
MiniCPM: Unveiling the potential of small language models with scalable training strategies
Hu, S., Tu, Y., Han, X., Cui, G., He, C., Zhao, W., Long, X., Zheng, Z., Fang, Y., Huang, Y., Zhang, X., Thai, Z. L., Wang, C., Yao, Y., Zhao, C., Zhou, J., Cai, J., Zhai, Z., Ding, N., Jia, C., Zeng, G., dahai li, Liu, Z., and Sun, M · 2024
Later among the works it cites.
Loss landscape characterization of neural networks without over-parametrization
Islamov, R., Ajroldi, N., Orvieto, A., and Lucchi, A · 2024
Later among the works it cites.
Last iterate of SGD converges (even in unbounded domains), 2020
Orabona, F · 2024
Later among the works it cites.
Power scheduler: A batch size and token number agnostic learning rate scheduler
Shen, Y., Stallone, M., Mishra, M., Zhang, G., Tan, S., Prasad, A., Soria, A. M., Cox, D. D., and Panda, R · 2024
Later among the works it cites.
Small-scale proxies for large-scale transformer training instabilities
Wortsman, M., Liu, P. J., Xiao, L., Everett, K. E., Alemi, A. A., Adlam, B., Co-Reyes, J. D., Gur, I., Kumar, A., Novak, R., Pennington, J., Sohl-Dickstein, J., Xu, K., Lee, J., Gilmer, J., and Kornblith, S · 2024
Later among the works it cites.
Rethinking conventional wisdom in machine learning: From generalization to scaling
Xiao, L · 2024
Later among the works it cites.
No more Adam: Learning rate scaling at initialization is all you need
Xu, M., Xiang, L., Cai, X., and Wen, H · 2024
Later among the works it cites.
Understanding warmup-stable-decay learning rates: A river valley loss landscape view
Wen, K., Li, Z., Wang, J. S., Hall, D. L. W., Liang, P., and Ma, T · 2025
Closest in time.
Deconstructing what makes a good optimizer for autoregressive language models
Zhao, R., Morwani, D., Brandfonbrener, D., Vyas, N., and Kakade, S. M · 2025
Closest in time.