Incorporating external knowledge through pre-training for natural language to code generation
Frank F. Xu, Zhengbao Jiang, Pengcheng Yin, Bogdan Vasilescu, and Graham Neubig. 2020 · 2020
Later among the works it cites.
Program synthesis with large language models
Original
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021 · 2021
Later among the works it cites.
Evaluating large language models trained on code
Original
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Later among the works it cites.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, et al. 2021 · 2021
Later among the works it cites.
Measuring coding challenge competence with apps
Original
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al. 2021 · 2021
Later among the works it cites.
MTOP: A comprehensive multilingual task-oriented semantic parsing benchmark
Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad. 2021 · 2021
Later among the works it cites.
Lyra: A benchmark for turducken-style code generation
Original
Qingyuan Liang, Zeyu Sun, Qihao Zhu, Wenjie Zhang, Lian Yu, Yingfei Xiong, and Lu Zhang. 2021 · 2021
Later among the works it cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Original
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al. 2021 · 2021
Later among the works it cites.
Code generation from natural language with less prior knowledge and more monolingual data
Sajad Norouzi, Keyi Tang, and Yanshuai Cao. 2021 · 2021
Later among the works it cites.
An empirical cybersecurity evaluation of github copilot’s code contributions
Original
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri. 2021 · 2021
Later among the works it cites.
Multi-domain multilingual question answering
Sebastian Ruder and Avi Sil. 2021 · 2021
Later among the works it cites.
Zero-shot cross-lingual semantic parsing
Original
Tom Sherborne and Mirella Lapata. 2021 · 2021
Later among the works it cites.
Cross-lingual training with dense retrieval for document retrieval
Original
Peng Shi, Rui Zhang, He Bai, and Jimmy Lin. 2021 · 2021
Later among the works it cites.
CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation
Yue Wang, Weishi Wang, Shafiq Joty, and Steven C.H. Hoi. 2021 · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
The flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022 · 2022
Closest in time.
Summarizing source code using a neural attention model
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2016 · 2083
Closest in time.
A convolutional attention network for extreme summarization of source code
Miltiadis Allamanis, Hao Peng, and Charles Sutton. 2016 · 2091
Closest in time.