ERNIE: Enhanced Representation through Knowledge Integration
Original
Yu Sun, Shuohuan Wang, Yukun Li, et al · 1904
Earlier work this paper cites.
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding
Original
Yu Sun, Shuohuan Wang, Yukun Li, et al · 1907
Earlier work this paper cites.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Original
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, et al · 1909
Earlier work this paper cites.
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Original
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 1910
Earlier work this paper cites.
Byte Pair Encoding: A Text Compression Scheme That Accelerates Pattern Matching
Yusuke Shibata, Takuya Kida, Shuichi Fukamachi, et al · 1999
Earlier work this paper cites.
On Layer Normalization in the Transformer Architecture
Original
Ruibin Xiong, Yunchang Yang, Di He, et al · 2002
Earlier work this paper cites.
Language Models are Few-Shot Learners
Original
Tom B. Brown, Benjamin Mann, Nick Ryder, et al · 2005
Earlier work this paper cites.
Keynote Address: .QL for Source Code Analysis. In Seventh IEEE International Working Conference on Source Code Analysis and Manipulation (SCAM 2007) . 3–16
Oege de Moor, Mathieu Verbaere, Elnar Hajiyev, et al · 2007
Earlier work this paper cites.
Layer Normalization
Original
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Original
Diederik P. Kingma and Jimmy Ba. 2017 · 2017
Earlier work this paper cites.
Pinpoint: fast and precise sparse value flow analysis for million lines of code. In Proceedings of the 39th ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI 2018, Philadelphia, PA, USA, June 18-22, 2018 , Jeffrey S. Foster and Dan Grossman (Eds.). ACM, 693–706
Qingkai Shi, Xiao Xiao, Rongxin Wu, et al · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Original
Leo Gao, Stella Biderman, Sid Black, et al · 2020
Earlier work this paper cites.
Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Association for Computational Linguistics, Online, 38–45
Thomas Wolf, Lysandre Debut, Victor Sanh, et al · 2020
Earlier work this paper cites.
Program Synthesis with Large Language Models
Original
Jacob Austin, Augustus Odena, Maxwell Nye, et al · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Earlier work this paper cites.