Efficient large scale language modeling with mixtures of experts
Original
Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, et al · 2021
Later among the works it cites.
Program synthesis with large language models
Original
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton · 2021
Later among the works it cites.
FairScale: A general purpose modular PyTorch library for high performance and large scale training
Mandeep Baines, Shruti Bhosale, Vittorio Caggiano, Naman Goyal, Siddharth Goyal, Myle Ott, Benjamin Lefaudeux, Vitaliy Liptchinsky, Mike Rabbat, Sam Sheiffer, Anjali Sridhar, and Min Xu · 2021
Later among the works it cites.
Learning type annotation: Is big data enough?
Kevin Jesse, Premkumar T. Devanbu, and Toufique Ahmed · 2021
Later among the works it cites.
CodeXGlue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Later among the works it cites.
DOBF: A deobfuscation pre-training objective for programming languages
Baptiste Roziere, Marie-Anne Lachaux, Marc Szafraniec, and Guillaume Lample · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki · 2021
Later among the works it cites.
CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi · 2021
Later among the works it cites.
Reflective decoding: Beyond unidirectional generation with off-the-shelf language models
Peter West, Ximing Lu, Ari Holtzman, Chandra Bhagavatula, Jena Hwang, and Yejin Choi · 2021
Later among the works it cites.
Break-it-fix-it: Unsupervised learning for program repair
Michihiro Yasunaga and Percy Liang · 2021
Later among the works it cites.
Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections
Ruiqi Zhong, Kristy Lee, Zheng Zhang, and Dan Klein · 2021
Later among the works it cites.
Efficient training of language models to fill in the middle
Original
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen · 2022
Closest in time.
GPT-NeoX-20B: An open-source autoregressive language model
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach · 2022
Closest in time.
PaLM: Scaling language modeling with pathways
Original
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
CodeFill: Multi-token code completion by jointly learning from structure and naming sequences
Maliheh Izadi, Roberta Gismondi, and Georgios Gousios · 2022
Closest in time.
Deduplicating training data mitigates privacy risks in language models
Original
Nikhil Kandpal, Eric Wallace, and Colin Raffel · 2022
Closest in time.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2022
Closest in time.
Competition-level code generation with AlphaCode
Original
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al · 2022
Closest in time.
A conversational paradigm for program synthesis
Original
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2022
Closest in time.
Training language models to follow instructions with human feedback
Original
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe · 2022
Closest in time.
LaMDA: Language models for dialog applications
Original
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Closest in time.
Natural Language Processing with Transformers
Lewis Tunstall, Leandro von Werra, and Thomas Wolf · 2022
Closest in time.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2022
Closest in time.
A systematic evaluation of large language models of code
Original
Frank F Xu, Uri Alon, Graham Neubig, and Vincent J Hellendoorn · 2022
Closest in time.