J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young et al. , “Scaling language models: Methods, analysis & insights from training gopher,” arXiv preprint arXiv:2112.11446 , 2021
Original
2021
Cited alongside, same era.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” CoRR , vol. abs/2101.03961, 2021
Original
2021
Cited alongside, same era.
Google, “More efficient in-context learning with glam,” https://ai.googleblog.com/2021/12/more-efficient-in-context-learning-with.html , 2021
2021
Cited alongside, same era.
A. Yang, J. Lin, R. Men, C. Zhou, L. Jiang, X. Jia, A. Wang, J. Zhang, J. Wang, Y. Li, D. Zhang, W. Lin, L. Qu, J. Zhou, and H. Yang, “M6-t: Exploring sparse expert models and beyond,” 2021. [Online]. Available: https://arxiv.org/abs/2105.15082
Original
2021
Cited alongside, same era.
Y. J. Kim, A. A. Awan, A. Muzio, A. F. Cruz-Salinas, L. Lu, A. Hendy, S. Rajbhandari, Y. He, and H. H. Awadalla, “Scalable and efficient moe training for multitask multilingual models,” CoRR , vol. abs/2109.10465, 2021. [Online]. Available: https://arxiv.org/abs/2109.10465
Original
2021
Cited alongside, same era.
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, A. Phanishayee, and M. Zaharia, “Efficient large-scale language model training on gpu clusters using megatron-lm,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , ser. SC ’21. New York, NY, USA: Association for Computing Machinery, 2021. [Online]. Available: https://doi.org/10.1145/3458817.3476209
2021
Cited alongside, same era.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” arXiv preprint arXiv:2101.03961 , 2021
Original
2021
Cited alongside, same era.
A. Ivanov, N. Dryden, T. Ben-Nun, S. Li, and T. Hoefler, “Data movement is all you need: A case study on optimizing transformers,” Proceedings of Machine Learning and Systems , vol. 3, pp. 711–732, 2021
2021
Cited alongside, same era.
J. Fang, Y. Yu, C. Zhao, and J. Zhou, “Turbotransformers: an efficient gpu serving system for transformer models,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2021, pp. 389–402
2021
Cited alongside, same era.
S. Chen, S. Huang, S. Pandey, B. Li, G. R. Gao, L. Zheng, C. Ding, and H. Liu, “E.T.: re-thinking self-attention for transformer models on gpus,” in SC ’21: The International Conference for High Performance Computing, Networking, Storage and Analysis, St. Louis, Missouri, USA, November 14 - 19, 2021 , B. R. de Supinski, M. W. Hall, and T. Gamblin, Eds. ACM, 2021, pp. 25:1–25:18
2021
Cited alongside, same era.
M. Dehghani, A. Arnab, L. Beyer, A. Vaswani, and Y. Tay, “The efficiency misnomer,” ArXiv , vol. abs/2110.12894, 2021
Original
2021
Cited alongside, same era.
S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, and Y. He, “Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , ser. SC ’21, 2021
2021
Cited alongside, same era.