Fetching the paper…
Reading the bibliography…
The scaling of the optimal AdamW weight decay hyperparameter with model and dataset size is critical as we seek to build larger models, but is poorly understood.
A downsampled variant of imagenet as an alternative to the cifar datasets
Chrabaszcz, P., Loshchilov, I., and Hutter, F · 2017
Earlier work this paper cites.
L2 regularization versus batch and weight normalization
Van Laarhoven, T · 2017
Earlier work this paper cites.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Balles, L. and Hennig, P · 2018
Earlier work this paper cites.
Jia, X., Song, S., He, W., Wang, Y., Rong, H., Zhou, F., Xie, L., Guo, Z., Yang, Y., Yu, L., et al · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
An empirical model of large-batch training
McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D · 2018
Earlier work this paper cites.
Theoretical analysis of auto rate-tuning by batch normalization
Arora, S., Li, Z., and Lyu, K · 2019
Earlier work this paper cites.
Autoaugment: Learning augmentation strategies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V · 2019
Earlier work this paper cites.
Openwebtext corpus
Gokaslan, A., Cohen, V., Pavlick, E., and Tellex, S · 2019
Earlier work this paper cites.
An exponential learning rate schedule for deep learning
Li, Z. and Arora, S · 2019
Earlier work this paper cites.
When does label smoothing help?
Müller, R., Kornblith, S., and Hinton, G. E · 2019
Earlier work this paper cites.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2019
Earlier work this paper cites.
S4l: Self-supervised semi-supervised learning
Zhai, X., Oliver, A., Kolesnikov, A., and Beyer, L · 2019
Earlier work this paper cites.
Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate
Li, Z., Lyu, K., and Arora, S · 2020
Cited alongside, same era.
Spherical motion dynamics: Learning dynamics of normalized neural network using sgd and weight decay
Wan, R., Zhu, Z., Zhang, X., and Sun, J · 2021
Cited alongside, same era.
Robust training of neural networks using scale invariant architectures
Li, Z., Bhojanapalli, S., Zaheer, M., Reddi, S., and Kumar, S · 2022
Cited alongside, same era.
On the sdes and scaling rules for adaptive gradient algorithms
Malladi, S., Lyu, K., Panigrahi, A., and Arora, S · 2022
Cited alongside, same era.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J · 2022
Cited alongside, same era.
Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit
Bordelon, B., Noci, L., Li, M. B., Hanin, B., and Pehlevan, C · 2024
Closest in time.
Why do we need weight decay in modern deep learning?
D’Angelo, F., Andriushchenko, M., Varre, A., and Flammarion, N · 2024
Closest in time.
Scaling exponents across parameterizations and optimizers
Everett, K. E., Xiao, L., Wortsman, M., Alemi, A. A., Novak, R., Liu, P. J., Gur, I., Sohl-Dickstein, J., Kaelbling, L. P., Lee, J., and Pennington, J · 2024
Closest in time.
A large-scale exploration of µ-transfer
Lingle, L · 2024
Closest in time.
Super consistency of neural network landscapes and learning rate transfer
Noci, L., Meterez, A., Hofmann, T., and Orvieto, A · 2024
Closest in time.
How to jointly tune learning rate and weight decay for AdamW
Schaipp, F · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
How to scale your ema
Busbridge, D., Ramapuram, J., Ablin, P., Likhomanenko, T., Dhekane, E. G., Suau Cuadros, X., and Webb, R · 2023
Cited alongside, same era.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Cited alongside, same era.
Rotational equilibrium: How weight decay balances learning across neural networks
Kosson, A., Messmer, B., and Jaggi, M · 2023
Cited alongside, same era.
Technical report for stablelm-3b-4e1t
Tow, J., Bellagente, M., Mahan, D., and Ruiz, C. R · 2023
Cited alongside, same era.
Scaling optimal lr across token horizons
Bjorck, J., Benhaim, A., Chaudhary, V., Wei, F., and Song, X · 2024
Cited alongside, same era.
u-mu p: The unit-scaled maximal update parametrization
Blake, C., Eichenberg, C., Dean, J., Balles, L., Prince, L. Y., Deiseroth, B., Cruz-Salinas, A. F., Luschi, C., Weinbach, S., and Orr, D · 2024
Cited alongside, same era.
Batch size invariant adam
Wang, X. and Aitchison, L · 2024
Closest in time.
Small-scale proxies for large-scale transformer training instabilities
Wortsman, M., Liu, P. J., Xiao, L., Everett, K. E., Alemi, A. A., Adlam, B., Co-Reyes, J. D., Gur, I., Kumar, A., Novak, R., Pennington, J., Sohl-Dickstein, J., Xu, K., Lee, J., Gilmer, J., and Kornblith, S · 2024
Closest in time.
Don’t be lazy: Completep enables compute-efficient deep transformers
Dey, N., Zhang, B. C., Noci, L., Li, M., Bordelon, B., Bergsma, S., Pehlevan, C., Hanin, B., and Hestness, J · 2025
Closest in time.
μ \mu nit scaling: Simple and scalable fp8 llm training
Narayan, S., Gupta, A., Paul, M., and Blalock, D · 2025
Closest in time.
Schaipp, F., Hägele, A., Taylor, A., Simsekli, U., and Bach, F · 2025
Closest in time.
How does critical batch size scale in pre-training?
Zhang, H., Morwani, D., Vyas, N., Wu, J., Zou, D., Ghai, U., Foster, D., and Kakade, S. M · 2025
Closest in time.