Fetching the paper…
Reading the bibliography…
The Lion optimizer has been a promising competitor with the AdamW for training large AI models, with advantages on memory, computation, and sample efficiency.
A direct adaptive method for faster backpropagation learning: The rprop algorithm
Riedmiller, M. and Braun, H · 1993
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Revisiting distributed synchronous sgd
Chen, J., Pan, X., Monga, R., Bengio, S., and Jozefowicz, R · 2016
Earlier work this paper cites.
Sparse communication for distributed gradient descent
Aji, A. F. and Heafield, K · 2017
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, W. J · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2017
Earlier work this paper cites.
Asynchronous stochastic gradient descent with delay compensation
Zheng, S., Meng, Q., Wang, T., Chen, W., Yu, N., Ma, Z.-M., and Liu, T.-Y · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Cited alongside, same era.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V · 2018
Cited alongside, same era.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Robustness to unbounded smoothness of generalized signsgd
Crawshaw, M., Liu, M., Orabona, F., Zhang, W., and Zhuang, Z · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Stingy sketch: a sketch framework for accurate and fast frequency estimation
Li, H., Chen, Q., Zhang, Y., Yang, T., and Cui, B · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Cited alongside, same era.
Openwebtext corpus
Gokaslan, A. and Cohen, V · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Socialiqa: Commonsense reasoning about social interactions
Sap, M., Rashkin, H., Chen, D., LeBras, R., and Choi, Y · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Cited alongside, same era.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Later among the works it cites.
Chainedfilter: Combining membership filters by chain rule
Li, H., Wang, L., Chen, Q., Ji, J., Wu, Y., Zhao, Y., Yang, T., and Akella, A · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Momentum ensures convergence of signsgd under weaker assumptions
Sun, T., Wang, Q., Li, D., and Wang, B · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Accelerating distributed deep learning using lossless homomorphic compression, 2024
Li, H., Xu, Y., Chen, J., Dwivedula, R., Wu, W., He, K., Akella, A., and Kim, D · 2024
Closest in time.
Asynchronous local-sgd training for language modeling
Liu, B., Chhaparia, R., Douillard, A., Kale, S., Rusu, A. A., Shen, J., Szlam, A., and Ranzato, M · 2024
Closest in time.