Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Later among the works it cites.
Demix layers: Disentangling domains for modular language modeling
Original
Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A Smith, and Luke Zettlemoyer · 2021
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Original
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Later among the works it cites.
Multilingual lama: Investigating knowledge in multilingual pretrained language models
Original
Nora Kassner, Philipp Dufter, and Hinrich Schütze · 2021
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Original
Ofir Press, Noah A Smith, and Mike Lewis · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Original
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2021
Later among the works it cites.
Mod-squad: Designing mixture of experts as modular multi-task learners
Original
Zitian Chen, Yikang Shen, Mingyu Ding, Zhenfang Chen, Hengshuang Zhao, Erik Learned-Miller, and Chuang Gan · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Original
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Sparsely activated mixture-of-experts are robust multi-task learners
Original
Shashank Gupta, Subhabrata Mukherjee, Krishan Subudhi, Eduardo Gonzalez, Damien Jose, Ahmed H Awadallah, and Jianfeng Gao · 2022
Later among the works it cites.
Internet-augmented language models through few-shot prompting for open-domain question answering
Original
Angeliki Lazaridou, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev · 2022
Later among the works it cites.
Few-shot learning with multilingual generative language models
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al · 2022
Later among the works it cites.
Codegen: An open large language model for code with multi-turn program synthesis
Original
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2022
Later among the works it cites.
mgpt: Few-shot learners go multilingual
Original
Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Vladislav Mikhailov, Anastasia Kozlova, and Tatiana Shavrina · 2022
Later among the works it cites.
A length-extrapolatable transformer
Original
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei · 2022
Later among the works it cites.
Natural language processing with transformers
Lewis Tunstall, Leandro Von Werra, and Thomas Wolf · 2022
Later among the works it cites.
Mixture of attention heads: Selecting attention heads per token
Xiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou, Wenge Rong, and Zhang Xiong · 2022
Later among the works it cites.
St-moe: Designing stable and transferable sparse expert models
Original
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Original
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Closest in time.
Sparsegpt: Massive language models can be accurately pruned in one-shot, 2023
Elias Frantar and Dan Alistarh · 2023
Closest in time.
On the effect of dropping layers of pre-trained transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov · 2023
Closest in time.
Llama: Open and efficient foundation language models
Original
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.