Fetching the paper…
Reading the bibliography…
Generative models such as large language models are extensively used as code copilots and for whole program generation.
Codebleu: a method for automatic evaluation of code synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020 · 2009
Earlier work this paper cites.
Analytical and experimental comparison of six algorithms for the vertex cover problem
François Delbot and Christian Laforest. 2010 · 2010
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021 · 2021
Earlier work this paper cites.
Vulrepair: a t5-based automated software vulnerability repair
Michael Fu, Chakkrit Tantithamthavorn, Trung Le, Van Nguyen, and Dinh Phung. 2022 · 2022
Earlier work this paper cites.
Factuality enhanced language models for open-ended text generation
Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale Fung, Mohammad Shoeybi, and Bryan Catanzaro. 2022 · 2022
Earlier work this paper cites.
Securityeval dataset: Mining vulnerability examples to evaluate machine learning-based code generation techniques
Mohammed Latif Siddiq and Joanna C. S. Santos. 2022 · 2022
Earlier work this paper cites.
A systematic evaluation of large language models of code
Frank F. Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn. 2022 · 2022
Earlier work this paper cites.
Multipl-e: A scalable and polyglot approach to benchmarking neural code generation
F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M. Yee, Y. Zi, C. Anderson, M. Q. Feldman, A. Guha, M. Greenberg, and A. Jangda. 2023 · 2023
Earlier work this paper cites.
Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation
Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, and Yiling Lou. 2023 · 2023
Earlier work this paper cites.
How secure is code generated by chatgpt?
Raphaël Khoury, Anderson R. Avila, Jacob Brunelle, and Baba Mamadou Camara. 2023 · 2023
Earlier work this paper cites.
Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023 · 2023
Cited alongside, same era.
Generate and pray: Using sallms to evaluate the security of llm generated code
Mohammed Latif Siddiq and Joanna C. S. Santos. 2023 · 2023
Cited alongside, same era.
Planning with large language models for code generation
Shun Zhang, Zhenfang Chen, Yikang Shen, Mingyu Ding, Joshua B. Tenenbaum, and Chuang Gan. 2023 · 2023
Cited alongside, same era.
Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x
Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Lei Shen, Zihan Wang, Andi Wang, Yang Li, Teng Su, Zhilin Yang, and Jie Tang. 2023 · 2023
Cited alongside, same era.
GitHub Copilot Subscriber Count
2024 · 2024
INSIDE: LLMs’ internal states retain the power of hallucination detection
Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye. 2024 · 2024
Closest in time.
Dola: Decoding by contrasting layers improves factuality in large language models
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He. 2024 · 2024
Closest in time.
Large language models for software engineering: A systematic literature review
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024 · 2024
Closest in time.
Code security vulnerability repair using reinforcement learning with large language models
Nafis Tanveer Islam, Mohammad Bahrami Karkevandi, and Peyman Najafirad. 2024 · 2024
Closest in time.
SWE-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Google Gemini
2024 · 2024
Cited alongside, same era.
Meta Code Llama
2024 · 2024
Cited alongside, same era.
Microsoft Copilot
2024 · 2024
Cited alongside, same era.
OpenAI ChatGPT
2024 · 2024
Cited alongside, same era.
Unsupervised evaluation of code llms with round-trip correctness
Miltiadis Allamanis, Sheena Panthaplackel, and Pengcheng Yin. 2024 · 2024
Cited alongside, same era.
Closest in time.
Exploring and evaluating hallucinations in llm-powered code generation
Fang Liu, Yang Liu, Lin Shi, Houkun Huang, Ruifeng Wang, Zhen Yang, Li Zhang, Zhongqi Li, and Yuchi Ma. 2024 · 2024
Closest in time.
Llms for science: Usage for code generation and data analysis
Mohamed Nejjar, Luca Zacharias, Fabian Stiehle, and Ingo Weber. 2024 · 2024
Closest in time.
Codehalu: Code hallucinations in llms driven by execution-based verification
Yuchen Tian, Weixiang Yan, Qian Yang, Qian Chen, Wen Wang, Ziyang Luo, and Lei Ma. 2024 · 2024
Closest in time.
Hallucination is inevitable: An innate limitation of large language models
Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024 · 2024
Closest in time.
ICE-score: Instructing large language models to evaluate code
Terry Yue Zhuo. 2024 · 2024
Closest in time.