Fetching the paper…
Reading the bibliography…
In this work, we question the necessity of adaptive gradient methods for training deep neural networks.
Bhardwaj, K., Li, G., and Marculescu, R · 1910
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Nesterov, Y. E · 1983
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
Doddington, G · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Banerjee, S. and Lavie, A · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Adadelta: An adaptive learning rate method, 2012
Zeiler, M. D · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks, 2014
Graves, A · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Incorporating Nesterov Momentum into Adam
Dozat, T · 2016
Earlier work this paper cites.
A downsampled variant of imagenet as an alternative to the cifar datasets, 2017
Chrabaszcz, P., Loshchilov, I., and Hutter, F · 2017
Earlier work this paper cites.
On weight initialization in deep neural networks
Kumar, S. K · 2017
Earlier work this paper cites.
The e2e dataset: New challenges for end-to-end generation
Novikova, J., Dušek, O., and Rieser, V · 2017
Cited alongside, same era.
Vaswani, A · 2017
Cited alongside, same era.
signsgd: Compressed optimisation for non-convex problems, 2018
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Later among the works it cites.
On the variance of the adaptive learning rate and beyond, 2021
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J · 2021
Later among the works it cites.
Early convolutions help transformers see better, 2021
Xiao, T., Singh, M., Mintun, E., Darrell, T., Dollár, P., and Girshick, R · 2021
Later among the works it cites.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lee, N., Ajanthan, T., and Torr, P. H · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Cited alongside, same era.
Openwebtext corpus
Gokaslan, A. and Cohen, V · 2019
Cited alongside, same era.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Pruning neural networks at initialization: Why are we missing the mark?
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Ghorbani, B., Suo, D., Cardoze, D., Dahl, G., Cohen, J., Gilmer, J., Agarwal, N., Krishnan, S., Medapati, S., and Nado, Z · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models, 2022
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers, 2022
Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L · 2022
Later among the works it cites.
The case for 4-bit precision: k-bit inference scaling laws
Dettmers, T. and Zettlemoyer, L · 2023
Later among the works it cites.
Improving robustness with adaptive weight decay, 2023
Ghiasi, A., Shafahi, A., and Ardekani, R · 2023
Later among the works it cites.
Kunstner, F., Chen, J., Lavington, J. W., and Schmidt, M · 2023
Later among the works it cites.
Balance is essence: Accelerating sparse training via adaptive gradient correction, 2023
Lei, B., Xu, D., Zhang, R., He, S., and Mallick, B. K · 2023
Later among the works it cites.
Prodigy: An expeditiously adaptive parameter-free learner
Mishchenko, K. and Defazio, A · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Exploiting network compressibility and topology in zero-cost nas
Xiang, L., Hunter, R., Xu, M., Dudziak, Ł., and Wen, H · 2023
Later among the works it cites.
Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models
Gu, Y., Wang, X., Wu, J. Z., Shi, Y., Chen, Y., Fan, Z., Xiao, W., Zhao, R., Chang, S., Wu, W., et al · 2024
Closest in time.
Convergence of adam under relaxed assumptions
Li, H., Rakhlin, A., and Jadbabaie, A · 2024
Closest in time.
Riemannian preconditioned lora for fine-tuning foundation models
Zhang, F. and Pilanci, M · 2024
Closest in time.