Fetching the paper…
Reading the bibliography…
Training vision or language models on large datasets can take days, if not weeks.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Earlier work this paper cites.
Ha, D., Dai, A., and Le, Q. V · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection
Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., and Belongie, S · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Regularizing and optimizing lstm language models
Merity, S., Keskar, N. S., and Socher, R · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2017
Earlier work this paper cites.
Importance sampling for minibatches
Csiba, D. and Richtárik, P · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D. P., and Wilson, A. G · 2018
Cited alongside, same era.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A. and Fleuret, F · 2018
Cited alongside, same era.
Lookahead optimizer: k steps forward, 1 step back
Zhang, M. R., Lucas, J., Ba, J., and Hinton, G. E · 2019
Later among the works it cites.
Beyer, L., Hénaff, O. J., Kolesnikov, A., Zhai, X., and Oord, A. v. d · 2020
Later among the works it cites.
The early phase of neural network training
Frankle, J., Schwab, D. J., and Morcos, A. S · 2020
Later among the works it cites.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho*, K., and Geras*, K · 2020
Later among the works it cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
Lakshminarayanan, C. and Szepesvari, C · 2018
Cited alongside, same era.
Troubling trends in machine learning scholarship
Lipton, Z. C. and Steinhardt, J · 2018
Cited alongside, same era.
Iterate averaging as regularization for stochastic gradient descent
Neu, G. and Rosasco, L · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L., Ma, C., et al · 2018
Cited alongside, same era.
Selection via proxy: Efficient data selection for deep learning
Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Cited alongside, same era.
Later among the works it cites.
Long-tailed classification by keeping the good and removing the bad momentum causal effect
Tang, K., Huang, J., and Zhang, H · 2020
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Later among the works it cites.
composer
Team, T. M. M · 2021
Later among the works it cites.
Stochastic weight averaging revisited
Guo, H., Jin, J., and Liu, B · 2022
Closest in time.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Closest in time.
Trainable weight averaging for fast convergence and better generalization, 2022
Li, T., Huang, Z., Tao, Q., Wu, Y., and Huang, X · 2022
Closest in time.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Mindermann, S., Brauner, J. M., Razzak, M. T., Sharma, M., Kirsch, A., Xu, W., Höltgen, B., Gomez, A. N., Morisot, A., Farquhar, S., et al · 2022
Closest in time.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al · 2022
Closest in time.