Fetching the paper…
Reading the bibliography…
We introduce MPLSandbox, an out-of-the-box multi-programming language sandbox designed to provide unified and comprehensive feedback from compiler and analysis tools for Large Language Models (LLMs).
Object code optimization
Edward S Lowry and Cleburne W Medlock. 1969 · 1969
Earlier work this paper cites.
Specification-based code generation. In Twenty-Third Annual Hawaii International Conference on System Sciences , Vol. 2. IEEE Computer Society, 165–173
S Antoy, Paola Forcheri, and Maria Teresa Molfino. 1990 · 1990
Earlier work this paper cites.
Automated software test data generation
Bogdan Korel. 1990 · 1990
Earlier work this paper cites.
A virtual machine introspection based architecture for intrusion detection.. In Ndss , Vol. 3. San Diega, CA, 191–206
Tal Garfinkel, Mendel Rosenblum, et al · 2003
Earlier work this paper cites.
Isolated program execution: An application transparent approach for executing untrusted programs. In 19th Annual Computer Security Applications Conference, 2003. Proceedings. IEEE, 182–191
Zhenkai Liang, VN Venkatakrishnan, and R Sekar. 2003 · 2003
Earlier work this paper cites.
Statistical analyses and reproducible research
Robert Gentleman and Duncan Temple Lang. 2007 · 2007
Earlier work this paper cites.
Automated whitebox fuzz testing.. In NDSS , Vol. 8. 151–166
Patrice Godefroid, Michael Y Levin, David A Molnar, et al · 2008
Earlier work this paper cites.
Pybox-a python sandbox
Markus Engelberth, Jan Göbel, Christian Schönbein, and Felix C Freiling. 2012 · 2012
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
How many of all bugs do we find? a study of static bug detectors. In Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering . 317–328
Andrew Habib and Michael Pradel. 2018 · 2018
Earlier work this paper cites.
Deep learning based code smell detection
Hui Liu, Jiahao Jin, Zhifeng Xu, Yanzhen Zou, Yifan Bu, and Lu Zhang. 2019 · 2019
Earlier work this paper cites.
The art, science, and engineering of fuzzing: A survey
Valentin JM Manès, HyungSeok Han, Choongwoo Han, Sang Kil Cha, Manuel Egele, Edward J Schwartz, and Maverick Woo. 2019 · 2019
Earlier work this paper cites.
A transformer-based approach for source code summarization
Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2020 · 2020
Earlier work this paper cites.
Software vulnerability detection using deep neural networks: a survey
Guanjun Lin, Sheng Wen, Qing-Long Han, Jun Zhang, and Yang Xiang. 2020 · 2020
Earlier work this paper cites.
Unsupervised translation of programming languages
Baptiste Roziere, Marie-Anne Lachaux, Lowik Chanussot, and Guillaume Lample. 2020 · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Measuring Coding Challenge Competence With APPS. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual , Joaquin Vanschoren and Sai-Kit Yeung (Eds.)
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Earlier work this paper cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Chu Hong Hoi. 2022 · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al · 2022
Earlier work this paper cites.
Code smells detection and visualization: a systematic literature review
José Pereira dos Reis, Fernando Brito e Abreu, Glauco de Figueiredo Carneiro, and Craig Anslow. 2022 · 2022
Earlier work this paper cites.
Execution-based evaluation for open-domain code generation
Zhiruo Wang, Shuyan Zhou, Daniel Fried, and Graham Neubig. 2022 · 2022
Earlier work this paper cites.
gpt-3.5-turbo
2023 · 2023
Earlier work this paper cites.
Vulnerability Detection and Monitoring Using LLM. In 2023 IEEE 9th International Women in Engineering (WIE) Conference on Electrical and Computer Engineering (WIECON-ECE) . IEEE, 309–314
Vishwanath Akuthota, Raghunandan Kasula, Sabiha Tasnim Sumona, Masud Mohiuddin, Md Tanzim Reza, and Md Mizanur Rahman. 2023 · 2023
Earlier work this paper cites.
SantaCoder: don’t reach for the stars!
Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Munoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, et al · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Earlier work this paper cites.
Triangulating python performance issues with SCALENE. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) . 51–64
Emery D Berger, Sam Stern, and Juan Altmayer Pizzorno. 2023 · 2023
Earlier work this paper cites.
Conversing with copilot: Exploring prompt engineering for solving cs1 problems using natural language. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 . 1136–1142
Paul Denny, Viraj Kumar, and Nasser Giacaman. 2023 · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules. In The Twelfth International Conference on Learning Representations
Hung Le, Hailin Chen, Amrita Saha, Akash Gokul, Doyen Sahoo, and Shafiq Joty. 2023 · 2023
Cited alongside, same era.
Reflection-tuning: Data recycling improves llm instruction-tuning
Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Heng Huang, Jiuxiang Gu, and Tianyi Zhou. 2023a · 2023
Cited alongside, same era.
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
DeepSeek-AI. 2024 · 2024
Closest in time.
What’s Wrong with Your Code Generated by Large Language Models? An Extensive Study
Shihan Dou, Haoxiang Jia, Shenxi Wu, Huiyuan Zheng, Weikang Zhou, Muling Wu, Mingxu Chai, Jessica Fan, Caishuang Huang, Yunbo Tao, et al · 2024
Closest in time.
StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
Shihan Dou, Yan Liu, Haoxiang Jia, Limao Xiong, Enyu Zhou, Junjie Shan, Caishuang Huang, Wei Shen, Xiaoran Fan, Zhiheng Xi, et al · 2024
Closest in time.
Mercury: An efficiency benchmark for llm code synthesis
Mingzhe Du, Anh Tuan Luu, Bin Ji, and See-Kiong Ng. 2024a · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taco: Topics in algorithmic code generation dataset
Rongao Li, Jie Fu, Bo-Wen Zhang, Tao Huang, Zhihong Sun, Chen Lyu, Guang Liu, Zhi Jin, and Ge Li. 2023b · 2023
Cited alongside, same era.
RLTF: Reinforcement Learning from Unit Test Feedback
Jiate Liu, Yiqin Zhu, Kaiwen Xiao, Qiang Fu, Xiao Han, Wei Yang, and Deheng Ye. 2023 · 2023
Cited alongside, same era.
GPT-4 Technical Report
OpenAI. 2023 · 2023
Cited alongside, same era.
Jiho Shin, Clark Tang, Tahmineh Mohati, Maleknaz Nayebi, Song Wang, and Hadi Hemmati. 2023 · 2023
Cited alongside, same era.
Execution-based code generation using deep reinforcement learning
Parshin Shojaee, Aneesh Jain, Sindhu Tipirneni, and Chandan K Reddy. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Beta-Coder: On Value-Based Deep Reinforcement Learning for Program Synthesis. In The Twelfth International Conference on Learning Representations
Zishun Yu, Yunzhe Tao, Liyu Chen, Tao Sun, and Hongxia Yang. 2023 · 2023
Cited alongside, same era.
Rrhf: Rank responses to align language models with human feedback without tears
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. 2023 · 2023
Cited alongside, same era.
Xiaohu Du, Ming Wen, Jiahao Zhu, Zifan Xie, Bin Ji, Huijun Liu, Xuanhua Shi, and Hai Jin. 2024b · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Search-Based LLMs for Code Optimization
Shuzheng Gao, Cuiyun Gao, Wenchao Gu, and Michael Lyu. 2024 · 2024
Closest in time.
TestART: Improving LLM-based Unit Test via Co-evolution of Automated Generation and Repair Iteration
Siqi Gu, Chunrong Fang, Quanjun Zhang, Fangyuan Tian, and Zhenyu Chen. 2024 · 2024
Closest in time.
DeepSeek-Coder: When the Large Language Model Meets Programming – The Rise of Code Intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024 · 2024
Closest in time.
Training LLMs to Better Self-Debug and Explain Code
Nan Jiang, Xiaopeng Li, Shiqi Wang, Qiang Zhou, Soneya Binta Hossain, Baishakhi Ray, Varun Kumar, Xiaofei Ma, and Anoop Deoras. 2024a · 2024
Closest in time.
From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
Haolin Jin, Linghan Huang, Haipeng Cai, Jun Yan, Bo Li, and Huaming Chen. 2024 · 2024
Closest in time.
Code Summarization without Direct Access to Code-Towards Exploring Federated LLMs for Software Engineering. In Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering . 100–109
Jahnavi Kumar and Sridhar Chimalakonda. 2024 · 2024
Closest in time.
GRACE: Empowering LLM-based software vulnerability detection with graph structure and in-context learning
Guilong Lu, Xiaolin Ju, Xiang Chen, Wenlong Pei, and Zhilong Cai. 2024 · 2024
Closest in time.
Bridging Gaps in LLM Code Translation: Reducing Errors with Call Graphs and Bridged Debuggers. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 2448–2449
Yang Luo, Richard Yu, Fajun Zhang, Ling Liang, and Yongqiang Xiong. 2024 · 2024
Closest in time.
SpecGen: Automated Generation of Formal Program Specifications via Large Language Models
Lezhi Ma, Shangqing Liu, Yi Li, Xiaofei Xie, and Lei Bu. 2024 · 2024
Closest in time.
Using an llm to help with code understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13
Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024 · 2024
Closest in time.
Lost in translation: A study of bugs introduced by large language models while translating code. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13
Rangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar, Lambert Pouguem Wassi, Michele Merler, Boris Sobolev, Raju Pavuluri, Saurabh Sinha, and Reyhaneh Jabbarvand. 2024 · 2024
Closest in time.
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
Houxing Ren, Mingjie Zhan, Zhongyuan Wu, Aojun Zhou, Junting Pan, and Hongsheng Li. 2024 · 2024
Closest in time.
Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLM
Gabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang, Xiaofei Ma, Murali Krishna Ramanathan, and Baishakhi Ray. 2024 · 2024
Closest in time.
Unraveling the Potential of Large Language Models in Code Translation: How Far Are We?
Qingxiao Tao, Tingrui Yu, Xiaodong Gu, and Beijun Shen. 2024 · 2024
Closest in time.
Qwen2.5: A Party of Foundation Models
Qwen Team. 2024 · 2024
Closest in time.
Enhancing Trust in LLM-Generated Code Summaries with Calibrated Confidence Scores
Yuvraj Virk, Premkumar Devanbu, and Toufique Ahmed. 2024 · 2024
Closest in time.
Python Symbolic Execution with LLM-powered Code Generation
Wenhan Wang, Kaibo Liu, An Ran Chen, Ge Li, Zhi Jin, Gang Huang, and Lei Ma. 2024 · 2024
Closest in time.
Fuzz4all: Universal fuzzing with large language models. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13
Chunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel, and Lingming Zhang. 2024 · 2024
Closest in time.
Large language models for cyber security: A systematic literature review
HanXiang Xu, ShenAo Wang, Ningke Li, Yanjie Zhao, Kai Chen, Kailong Wang, Yang Liu, Ting Yu, and HaoYu Wang. 2024 · 2024
Closest in time.
Rectifier: Code Translation with Corrector via LLMs
Xin Yin, Chao Ni, Tien N Nguyen, Shaohua Wang, and Xiaohu Yang. 2024 · 2024
Closest in time.
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Qihao Zhu, Daya Guo, Zhihong Shao, Dejian Yang, Peiyi Wang, Runxin Xu, Y Wu, Yukun Li, Huazuo Gao, Shirong Ma, et al · 2024
Closest in time.