M. Shoeybi et al. , “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism,” Mar. 2020, arXiv:1909.08053 [cs]. [Online]. Available: http://arxiv.org/abs/1909.08053
Original
1909
Earlier work this paper cites.
P. Gage, “A New Algorithm for Data Compression,” C Users Journal , vol. 12, no. 2, pp. 23–38, Feb. 1994
1994
Earlier work this paper cites.
R. Mihalcea, H. Liu, and H. Lieberman, “NLP (Natural Language Processing) for NLP (Natural Language Programming),” in Computational Linguistics and Intelligent Text Processing , A. Gelbukh, Ed. Springer Berlin Heidelberg, 2006, pp. 319–330
2006
Earlier work this paper cites.
C. B. Harris and I. G. Harris, “GLAsT: Learning formal grammars to translate natural language specifications into hardware assertions,” in Design, Automation Test in Europe Conf. Exhibition (DATE) , 2016, pp. 966–971
2016
Earlier work this paper cites.
A. Vaswani et al. , “Attention is All you Need,” in Advances in Neural Information Processing Systems 30 , I. Guyon et al. , Eds. Curran Associates, Inc., 2017, pp. 5998–6008
2017
Earlier work this paper cites.