Solving linear algebra by program synthesis
Iddo Drori and Nakul Verma. 2021 · 2021
Later among the works it cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2021 · 2021
Later among the works it cites.
Graphcodebert: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. 2021 · 2021
Later among the works it cites.
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Later among the works it cites.
Contrastive code representation learning
Paras Jain, Ajay Jain, Tianjun Zhang, Pieter Abbeel, Joseph Gonzalez, and Ion Stoica. 2021 · 2021
Later among the works it cites.
Deep graph matching and searching for semantic code retrieval
Xiang Ling, Lingfei Wu, Sai gang Wang, Gaoning Pan, Tengfei Ma, Fangli Xu, Alex X. Liu, Chunming Wu, and Shouling Ji. 2021 · 2021
Later among the works it cites.
Codexglue - a machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021 · 2021
Later among the works it cites.
Cotext: Multi-task learning with code-text transformer
Long N. Phan, Hieu Tran, Daniel Le, Hieu Nguyen, James T. Anibal, Alec Peltekian, and Yanfang Ye. 2021 · 2021
Later among the works it cites.
Project codenet: A large-scale ai for code dataset for learning a diversity of coding tasks
Original
Ruchi Puri, David S. Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladmir Zolotov, Julian Dolby, Jie Chen, Mihir R. Choudhury, Lindsey Decker, Veronika Thost, Luca Buratti, Saurabh Pujar, and Ulrich Finkler. 2021 · 2021
Later among the works it cites.
Fsf-funded call for white papers on philosophical and legal questions around copilot
Donald Robertson. 2021 · 2021
Later among the works it cites.
Dobf: A deobfuscation pre-training objective for programming languages
Baptiste Rozière, Marie-Anne Lachaux, Marc Szafraniec, and Guillaume Lample. 2021 · 2021
Later among the works it cites.
Solving probability and statistics problems by program synthesis
Leonard Tang, Elizabeth Ke, Nikhil Singh, Nakul Verma, and Iddo Drori. 2021 · 2021
Later among the works it cites.
Mulcode: A multi-task learning approach for source code understanding
Deze Wang, Yue Yu, Shanshan Li, Wei Dong, Ji Wang, and Qing Liao. 2021a · 2021
Later among the works it cites.
D2a: A dataset built for ai-based vulnerability detection methods using differential analysis
Yunhui Zheng, Saurabh Pujar, Burn Lewis, Luca Buratti, Edward Epstein, Bo Yang, Jim Laredo, Alessandro Morari, and Zhonglai Su. 2021 · 2021
Later among the works it cites.
The robots are coming: Exploring the implications of openai codex on introductory programming
James Finnie-Ansley, Paul Denny, Brett A. Becker, Andrew Luxton-Reilly, and James Prather. 2022 · 2022
Closest in time.
Unixcoder: Unified cross-modal pre-training for code representation
Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022 · 2022
Closest in time.
The stack: 3 tb of permissively licensed source code
Original
Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries. 2022 · 2022
Closest in time.
Competition-level code generation with alphacode
Yujia Li, David H. Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals. 2022 · 2022
Closest in time.
Automatically generating cs learning materials with large language models
Stephen MacNeil, Andrew Tran, Juho Leinonen, Paul Denny, Joanne Kim, Arto Hellas, Seth Bernstein, and Sami Sarsa. 2022 · 2022
Closest in time.
Formal mathematics statement curriculum learning
Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, and Ilya Sutskever. 2022 · 2022
Closest in time.
Progprompt: Generating situated robot task plans using large language models
Original
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. 2022 · 2022
Closest in time.
Natural Language Processing with Transformers: Building Language Applications with Hugging Face
L. Tunstall, L. von Werra, and T. Wolf. 2022 · 2022
Closest in time.
A systematic evaluation of large language models of code
Original
Frank F. Xu, Uri Alon, Graham Neubig, and Vincent J. Hellendoorn. 2022 · 2022
Closest in time.