Fetching the paper…
Reading the bibliography…
AI-driven program repair uses AI models to repair buggy software by producing patches.
Defects4j: A database of existing faults to enable controlled testing studies for java programs
René Just, Darioush Jalali, and Michael D Ernst · 2014
Earlier work this paper cites.
The manybugs and introclass benchmarks for automated repair of c programs
Claire Le Goues, Neal Holtschulte, Edward K Smith, Yuriy Brun, Premkumar Devanbu, Stephanie Forrest, and Westley Weimer · 2015
Earlier work this paper cites.
Bugs. jar: A large-scale, diverse dataset of real-world java bugs
Ripon K Saha, Yingjun Lyu, Wing Lam, Hiroaki Yoshida, and Mukul R Prasad · 2018
Earlier work this paper cites.
Bugsjs: a benchmark of javascript bugs
Péter Gyimesi, Béla Vancsics, Andrea Stocco, Davood Mazinanian, Arpád Beszédes, Rudolf Ferenc, and Ali Mesbah · 2019
Earlier work this paper cites.
Bears: An extensible java bug benchmark for automatic program repair studies
Fernanda Madeiral, Simon Urli, Marcelo Maia, and Martin Monperrus · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Fixjs: a dataset of bug-fixing javascript commits
Viktor Csuvik and László Vidács · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Earlier work this paper cites.
Less training, more repairing please: revisiting automated program repair via zero-shot learning
Chunqiu Steven Xia and Lingming Zhang · 2022
Earlier work this paper cites.
A systematic evaluation of large language models of code
Frank F Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn · 2022
Cited alongside, same era.
Minecraft: Automated mining of software bug fixes with precise code context
Sai Krishna Avula, Venkatesh Vobbilisetti, and Shouvick Mondal · 2023
Cited alongside, same era.
Automated repair of programs from large language models
Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan · 2023
Cited alongside, same era.
Impact of code language models on automated program repair
Nan Jiang, Kevin Liu, Thibaud Lutellier, and Lin Tan · 2023
Cited alongside, same era.
Llm is like a box of chocolates: the non-determinism of chatgpt in code generation
Shuyin Ouyang, Jie M Zhang, Mark Harman, and Meng Wang · 2023
Cited alongside, same era.
Program repair competition
Cigar: Cost-efficient program repair with llms
Dávid Hidvégi, Khashayar Etemadi, Sofia Bobadilla, and Martin Monperrus · 2024
Closest in time.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica · 2024
Closest in time.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan · 2024
Closest in time.
Xcodeeval: An execution-based large scale multilingual multitask benchmark for code understanding, generation, translation and retrieval
Mohammad Abdullah Matin Khan, M Saiful Bari, Do Long, Weishi Wang, Md Rizwan Parvez, and Shafiq Joty · 2024
Closest in time.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ridwan Shariffdeen, Martin Mirchev, and Abhik Roychoudhury · 2023
Cited alongside, same era.
A survey of learning-based automated program repair
Quanjun Zhang, Chunrong Fang, Yuxiang Ma, Weisong Sun, and Zhenyu Chen · 2023
Cited alongside, same era.
On the reproducibility of software defect datasets
Hao-Nan Zhu and Cindy Rubio-González · 2023
Cited alongside, same era.
Claude 3.5 sonnet, June 2024
Anthropic · 2024
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference, 2024
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
Cited alongside, same era.
Yihong Dong, Xue Jiang, Huanyu Liu, Zhi Jin, and Ge Li · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Cited alongside, same era.
Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, et al · 2024
Closest in time.
On leakage of code generation evaluation datasets
Alexandre Matton, Tom Sherborne, Dennis Aumiller, Elena Tommasone, Milad Alizadeh, Jingyi He, Raymond Ma, Maxime Voisin, Ellen Gilsenan-McMahon, and Matthias Gallé · 2024
Closest in time.
Large enough, July 2024
Mistral · 2024
Closest in time.
The fact selection problem in llm-based program repair
Nikhil Parasaram, Huijie Yan, Boyu Yang, Zineb Flahy, Abriele Qudsi, Damian Ziaber, Earl Barr, and Sergey Mechtaev · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Gitbug-java: A reproducible benchmark of recent java bugs
André Silva, Nuno Saavedra, and Martin Monperrus · 2024
Closest in time.
Qwen2.5: A party of foundation models, September 2024
Qwen Team · 2024
Closest in time.
Agentless: Demystifying llm-based software engineering agents
Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang · 2024
Closest in time.