Fetching the paper…
Reading the bibliography…
This paper provides a comprehensive review of the current methods and metrics used to evaluate the performance of Large Language Models (LLMs) in code generation tasks.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Defects4j: a database of existing faults to enable controlled testing studies for java programs
René Just, Darioush Jalali, and Michael D. Ernst · 2014
Earlier work this paper cites.
Convolutional neural networks over tree structures for programming language processing
Lili Mou, Ge Li, Lu Zhang, Tao Wang, and Zhi Jin · 2016
Earlier work this paper cites.
Mapping language to code in programmatic context
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 2019
Earlier work this paper cites.
Learning gaussian policies from corrective human feedback
Daan Wout, Jan Scholten, Carlos Celemin, and Jens Kober · 2019
Earlier work this paper cites.
Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task, 2019
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev · 2019
Earlier work this paper cites.
Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks
Yaqin Zhou, Shangqing Liu, Jing Kai Siow, Xiaoning Du, and Yang Liu · 2019
Earlier work this paper cites.
Human or machine: Automating human likeliness evaluation of nlg texts
Erion Çano and Ondřej Bojar · 2020
Earlier work this paper cites.
Codebleu: a method for automatic evaluation of code synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps, 2021
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Semantic-aware binary code representation with bert. arxiv 2021
H Koo, S Park, D Choi, and T Kim · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Earlier work this paper cites.
Codenet: A large-scale ai for code dataset for learning a diversity of coding tasks, 2021
Ruchir Puri, David S. Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladimir Zolotov, Julian Dolby, Jie Chen, Mihir Choudhury, Lindsey Decker, Veronika Thost, Luca Buratti, Saurabh Pujar, Shyam Ramji, Ulrich Finkler, Susan Malaika, and Frederick Reiss · 2021
Earlier work this paper cites.
Multipl-e: A scalable and extensible approach to benchmarking neural code generation
Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, et al · 2022
Earlier work this paper cites.
Semantic similarity metrics for evaluating source code summarization
Sakib Haque, Zachary Eberhart, Aakash Bansal, and Collin McMillan · 2022
Earlier work this paper cites.
Interactive code generation via test-driven user-intent formalization
Shuvendu K Lahiri, Aaditya Naik, Georgios Sakkas, Piali Choudhury, Curtis von Veh, Madanlal Musuvathi, Jeevana Priya Inala, Chenglong Wang, and Jianfeng Gao · 2022
Cited alongside, same era.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, and Schrittwieser · 2022
Cited alongside, same era.
Compilable neural code generation with compiler feedback
Xin Wang, Yasheng Wang, Yao Wan, Fei Mi, Yitong Li, Pingyi Zhou, Jin Liu, Hao Wu, Xin Jiang, and Qun Liu · 2022
Cited alongside, same era.
Xlcost: A benchmark dataset for cross-lingual code intelligence
Ming Zhu, Aneesh Jain, Karthik Suresh, Roshan Ravindran, Sindhu Tipirneni, and Chandan K Reddy · 2022
Cited alongside, same era.
Website, 2023
Levenshtein distance · 2023
Enhancing large language models for secure code generation: A dataset-driven study on vulnerability mitigation
Jiexin Wang, Liuwen Cao, Xitong Luo, Zhiping Zhou, Jiayuan Xie, Adam Jatowt, and Yi Cai · 2023
Later among the works it cites.
Leti: Learning to generate from textual interactions
Xingyao Wang, Hao Peng, Reyhaneh Jabbarvand, and Heng Ji · 2023
Later among the works it cites.
Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization
Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, et al · 2023
Later among the works it cites.
Codescope: An execution-based multilingual multitask multidimensional benchmark for evaluating llms on code understanding and generation
Weixiang Yan, Haitian Liu, Yunkun Wang, Yunzhe Li, Qian Chen, Wen Wang, Tingyu Lin, Weishan Zhao, Li Zhu, Shuiguang Deng, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring large language models for code explanation
Paheli Bhattacharya, Manojit Chakraborty, Kartheek NSN Palepu, Vikas Pandey, Ishan Dindorkar, Rakesh Rajpurohit, and Rishabh Gupta · 2023
Cited alongside, same era.
Can it edit? evaluating the ability of large language models to follow code editing instructions
Federico Cassano, Luisa Li, Akul Sethi, Noah Shinn, Abby Brennan-Jones, Anton Lozhkov, Carolyn Anderson, and Arjun Guha · 2023
Cited alongside, same era.
Aligning offline metrics and human judgments of value for code generation models
Victor Dibia, Adam Fourney, Gagan Bansal, Forough Poursabzi-Sangdeh, Han Liu, and Saleema Amershi · 2023
Cited alongside, same era.
Large language models for software engineering: Survey and open problems
Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M Zhang · 2023
Cited alongside, same era.
Fixeval: Execution-based evaluation of program fixes for programming problems, 2023
Md Mahim Anjum Haque, Wasi Uddin Ahmad, Ismini Lourentzou, and Chris Brown · 2023
Cited alongside, same era.
Verilogeval: Evaluating large language models for verilog code generation
Mingjie Liu, Nathaniel Pinckney, Brucek Khailany, and Haoxing Ren · 2023
Cited alongside, same era.
Recent advances in natural language processing via large pre-trained language models: A survey
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth · 2023
Cited alongside, same era.
Asteria-pro: Enhancing deep learning-based binary code similarity detection by incorporating domain knowledge
Shouguo Yang, Chaopeng Dong, Yang Xiao, Yiran Cheng, Zhiqiang Shi, Zhi Li, and Limin Sun · 2023
Later among the works it cites.
Codereval: A benchmark of pragmatic code generation with generative pre-trained models
Hao Yu, Bo Shen, Dezhi Ran, Jiaxin Zhang, Qi Zhang, Yuchi Ma, Guangtai Liang, Ying Li, Qianxiang Wang, and Tao Xie · 2023
Later among the works it cites.
Codegen-test: An automatic code generation model integrating program test information
Maosheng Zhong, Zhixiang Wang, Gen Liu, Youde Chen, Huizhu Liu, and Ruping Wu · 2023
Later among the works it cites.
Copilot evaluation harness: Evaluating llm-guided software programming
Anisha Agarwal, Aaron Chan, Shubham Chandel, Jinu Jang, Shaun Miller, Roshanak Zilouchian Moghaddam, Yevhen Mohylevskyy, Neel Sundaresan, and Michele Tufano · 2024
Closest in time.
Automating the correctness assessment of ai-generated code for security contexts
Domenico Cotroneo, Alessio Foggia, Cristina Improta, Pietro Liguori, and Roberto Natella · 2024
Closest in time.
Mercury: An efficiency benchmark for llm code synthesis
Mingzhe Du, Anh Tuan Luu, Bin Ji, and See-Kiong Ng · 2024
Closest in time.
Effibench: Benchmarking the efficiency of automatically generated code
Dong Huang, Jie M Zhang, Yuhao Qing, and Heming Cui · 2024
Closest in time.
Devbench: A comprehensive benchmark for software development, 2024
Bowen Li, Wenhan Wu, Ziwei Tang, Lin Shi, John Yang, Jinyang Li, Shunyu Yao, Chen Qian, Binyuan Hui, Qicheng Zhang, Zhiyin Yu, He Du, Ping Yang, Dahua Lin, Chao Peng, and Kai Chen · 2024
Closest in time.
No need to lift a finger anymore? assessing the quality of code generation by chatgpt
Zhijie Liu, Yutian Tang, Xiapu Luo, Yuming Zhou, and Liang Feng Zhang · 2024
Closest in time.
The realhumaneval: Evaluating large language models’ abilities to support programmers
Hussein Mozannar, Valerie Chen, Mohammed Alsobay, Subhro Das, Sebastian Zhao, Dennis Wei, Manish Nagireddy, Prasanna Sattigeri, Ameet Talwalkar, and David Sontag · 2024
Closest in time.
Autosurvey: Large language models can automatically write surveys
Yidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang, Xin Zhang, Zhen Wu, Meishan Zhang, Xinyu Dai, Min Zhang, Qingsong Wen, et al · 2024
Closest in time.
Plot2code: A comprehensive benchmark for evaluating multi-modal large language models in code generation from scientific plots, 2024
Chengyue Wu, Yixiao Ge, Qiushan Guo, Jiahao Wang, Zhixuan Liang, Zeyu Lu, Ying Shan, and Ping Luo · 2024
Closest in time.
Coderujb: An executable and unified java benchmark for practical programming scenarios
Zhengran Zeng, Yidong Wang, Rui Xie, Wei Ye, and Shikun Zhang · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.