Fetching the paper…
Reading the bibliography…
Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities.
A complexity measure
Thomas J. McCabe · 1976
Earlier work this paper cites.
Elements of Software Science (Operating and programming systems series)
Maurice H. Halstead · 1977
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Defects4J: A Database of existing faults to enable controlled testing studies for Java programs
René Just, Darioush Jalali, and Michael D. Ernst · 2014
Earlier work this paper cites.
Mining the modern code review repositories: A dataset of people, process and product
Xin Yang, Raula Gaikovina Kula, Norihiro Yoshida, and Hajimu Iida · 2016
Earlier work this paper cites.
Learning to represent programs with graphs
Miltiadis Allamanis, Marc Brockschmidt, and Mahmoud Khademi · 2017
Earlier work this paper cites.
Deepfix: Fixing common c language errors by deep learning
Rahul Gupta, Soham Pal, Aditya Kanade, and Shirish Shevade · 2017
Earlier work this paper cites.
Entropy guided spectrum based bug localization using statistical language model
Saikat Chakraborty, Yujian Li, Matt Irvine, Ripon Saha, and Baishakhi Ray · 2018
Earlier work this paper cites.
Shaping program repair space with existing patches and similar code
Jiajun Jiang, Yingfei Xiong, Hongyu Zhang, Qing Gao, and Xiangqun Chen · 2018
Earlier work this paper cites.
Automatic software repair
Martin Monperrus · 2018
Earlier work this paper cites.
Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev · 2018
Earlier work this paper cites.
Automated program repair
Claire Le Goues, Michael Pradel, and Abhik Roychoudhury · 2019
Earlier work this paper cites.
How often do single-statement bugs occur? the manysstubs4j dataset
Rafael-Michael Karampatsis and Charles Sutton · 2019
Earlier work this paper cites.
Precise learn-to-rank fault localization using dynamic and static features of target programs
Yunho Kim, Seokhyeon Mun, Shin Yoo, and Moonzoo Kim · 2019
Earlier work this paper cites.
Language tasks and language games: On methodology in current natural language processing research, 2019
David Schlangen · 2019
Earlier work this paper cites.
Program synthesis with large language models, 2021
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton · 2021
Earlier work this paper cites.
What will it take to fix benchmarking in natural language understanding?, 2021
Samuel R. Bowman and George E. Dahl · 2021
Earlier work this paper cites.
On multi-modal learning of editing source code, 2021
Saikat Chakraborty and Baishakhi Ray · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, and Jared Kaplan et. al · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps, 2021
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Commitbert: Commit message generation using pre-trained programming language model, 2021
Tae-Hwan Jung · 2021
Earlier work this paper cites.
Dynabench: Rethinking benchmarking in nlp, 2021
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, and Zhengxuan Wu et. al · 2021
Cited alongside, same era.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, and Ambrosio Blanco et. al · 2021
Cited alongside, same era.
Research community dynamics behind popular ai benchmarks
Fernando Martínez-Plumed, Pablo Barredo, Seán Ó hÉigeartaigh, and José Hernández-Orallo · 2021
Cited alongside, same era.
Towards automating code review activities, 2021
Rosalia Tufano, Luca Pascarella, Michele Tufano, Denys Poshyvanyk, and Gabriele Bavota · 2021
Cited alongside, same era.
Multipl-e: A scalable and extensible approach to benchmarking neural code generation, 2022
Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda · 2022
Cited alongside, same era.
Automated repair of programs from large language models, 2023
Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan · 2023
Closest in time.
Incoder: A generative model for code infilling and synthesis, 2023
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Wen tau Yih, Luke Zettlemoyer, and Mike Lewis · 2023
Closest in time.
Ai safety subproblems for software engineering researchers, 2023
David Gros, Prem Devanbu, and Zhou Yu · 2023
Closest in time.
Large language models for software engineering: A systematic literature review, 2023
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang · 2023
Closest in time.
Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models, 2023
Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Leon Song, Samyam Rajbhandari, and Yuxiong He · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
Program repair, 2022
Xiang Gao, Yannic Noller, and Abhik Roychoudhury · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
Ds-1000: A natural and reliable benchmark for data science code generation, 2022
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Scott Wen tau Yih, Daniel Fried, Sida Wang, and Tao Yu · 2022
Cited alongside, same era.
Holistic evaluation of language models, 2022
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, and Michihiro Yasunaga et. al · 2022
Cited alongside, same era.
Mapping global dynamics of benchmark creation and saturation in artificial intelligence
Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec, Janina Brauner, and Matthias Samwald · 2022
Cited alongside, same era.
Less training, more repairing please: revisiting automated program repair via zero-shot learning
Chunqiu Steven Xia and Lingming Zhang · 2022
Cited alongside, same era.
Large language models are few-shot testers: Exploring llm-based general bug reproduction, 2023
Sungmin Kang, Juyeon Yoon, and Shin Yoo · 2023
Closest in time.
Large sequence models for software development activities, 2023
Petros Maniatis, Daniel Tarlow, and Google DeepMind · 2023
Closest in time.
Better automatic program repair by using bug reports and tests together, 2023
Manish Motwani and Yuriy Brun · 2023
Closest in time.
Octopack: Instruction tuning code large language models, 2023
Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, and Shayne Longpre · 2023
Closest in time.
Measuring the impact of programming language distribution, 2023
Gabriel Orlanski, Kefan Xiao, Xavier Garcia, Jeffrey Hui, Joshua Howland, Jonathan Malmaud, Jacob Austin, Rishabh Singh, and Michele Catasta · 2023
Closest in time.
Code llama: Open foundation models for code, 2023
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, and Xiaoqing Ellen Tan et. al · 2023
Closest in time.
An analysis of the automatic bug fixing performance of chatgpt, 2023
Dominik Sobania, Martin Briesch, Carol Hanna, and Justyna Petke · 2023
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, and Abu Awal Md Shoeb et. al · 2023
Closest in time.
Software testing with large language model: Survey, landscape, and vision, 2023
Junjie Wang, Yuchao Huang, Chunyang Chen, Zhe Liu, Song Wang, and Qing Wang · 2023
Closest in time.
Conversational automated program repair, 2023
Chunqiu Steven Xia and Lingming Zhang · 2023
Closest in time.
Universal fuzzing via large language models
Chunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel, and Lingming Zhang · 2023
Closest in time.
Intercode: Standardizing and benchmarking interactive coding with execution feedback, 2023
John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao · 2023
Closest in time.
Codereval: A benchmark of pragmatic code generation with generative pre-trained models, 2023
Hao Yu, Bo Shen, Dezhi Ran, Jiaxin Zhang, Qi Zhang, Yuchi Ma, Guangtai Liang, Ying Li, Tao Xie, and Qianxiang Wang · 2023
Closest in time.
Large language models meet nl2code: A survey, 2023
Daoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Yongji Wang, and Jian-Guang Lou · 2023
Closest in time.
Towards an understanding of large language models in software engineering tasks
Zibin Zheng, Kaiwen Ning, Jiachi Chen, Yanlin Wang, Wenqing Chen, Lianghong Guo, and Weicheng Wang · 2023
Closest in time.
Webarena: A realistic web environment for building autonomous agents, 2023
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig · 2023
Closest in time.