Fetching the paper…
Reading the bibliography…
With the rapidly increasing capabilities and adoption of code agents for AI-assisted coding, safety concerns, such as generating or executing risky code, have become significant barriers to the real-world deployment of these agents.
Docker: lightweight linux containers for consistent development and deployment
D. Merkel et al · 2014
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Earlier work this paper cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
H. Pearce, B. Ahmad, B. Tan, B. Dolan-Gavitt, and R. Karri · 2022
Earlier work this paper cites.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, B. Hui, L. Ji, M. Li, J. Lin, R. Lin, D. Liu, G. Liu, C. Lu, K. Lu, J. Ma, R. Men, X. Ren, X. Ren, C. Tan, S. Tan, J. Tu, P. Wang, S. Wang, W. Wang, S. Wu, B. Xu, J. Xu, A. Yang, H. Yang, J. Yang, S. Yang, Y. Yao, B. Yu, H. Yuan, Z. Yuan, J. Zhang, X. Zhang, Y. Zhang, Z. Zhang, C. Zhou, J. Zhou, X. Zhou, and T. Zhu · 2023
Earlier work this paper cites.
Purple llama cyberseceval: A secure coding benchmark for language models
M. Bhatt, S. Chennabasappa, C. Nikolaidis, S. Wan, I. Evtimov, D. Gabi, D. Song, F. Ahmad, C. Aschermann, L. Fontana, et al · 2023
Earlier work this paper cites.
Diagnosing and fixing memory leaks in python, 2023
GeeksforGeeks · 2023
Earlier work this paper cites.
Agentbench: Evaluating llms as agents
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, et al · 2023
Earlier work this paper cites.
Testing language model agents safely in the wild
S. Naihin, D. Atkinson, M. Green, M. Hamadi, C. Swift, D. Schonholtz, A. T. Kalai, and D. Bau · 2023
Earlier work this paper cites.
An attacker’s dream? exploring the capabilities of chatgpt for developing malware
Y. M. P. Pa, S. Tanizaki, T. Kou, M. van Eeten, K. Yoshioka, and T. Matsumoto · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson · 2023
Earlier work this paper cites.
Code llama: Open foundation models for code
B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, T. Remez, J. Rapin, et al · 2023
Cited alongside, same era.
Reflexion: language agents with verbal reinforcement learning
N. Shinn, F. Cassano, B. Labash, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson · 2023
Cited alongside, same era.
https://www.virustotal.com/gui/home/upload
Virustotal · 2024
Cited alongside, same era.
Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models
CodeLMSec benchmark: Systematically evaluating and finding security vulnerabilities in black-box code language models
H. Hajipour, K. Hassler, T. Holz, L. Schönherr, and M. Fritz · 2024
Closest in time.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
L. Li, B. Dong, R. Wang, X. Hu, W. Zuo, D. Lin, Y. Qiao, and J. Shao · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks · 2024
Closest in time.
Common weakness enumeration (cwe) list version 4.14, a community-developed dictionary of software weaknesses types
T. M. C. (MITRE) · 2024
Closest in time.
Identifying the risks of lm agents with an lm-emulated sandbox
Y. Ruan, H. Dong, A. Wang, S. Pitis, Y. Zhou, J. Ba, Y. Dubois, C. J. Maddison, and T. Hashimoto · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Bhatt, S. Chennabasappa, Y. Li, C. Nikolaidis, D. Song, S. Wan, F. Ahmad, C. Aschermann, Y. Chen, D. Kapil, et al · 2024
Cited alongside, same era.
Teaching large language models to self-debug
X. Chen, M. Lin, N. Schärli, and D. Zhou · 2024
Cited alongside, same era.
Agentdojo: A dynamic environment to evaluate attacks and defenses for llm agents
E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer, and F. Tramèr · 2024
Cited alongside, same era.
Introducing devin, the first ai software engineer
Devin · 2024
Cited alongside, same era.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. Li, et al · 2024
Cited alongside, same era.
Craft: Customizing llms by creating and retrieving from specialized toolsets
L. Yuan, Y. Chen, X. Wang, Y. R. Fung, H. Peng, and H. Ji
Cited in the paper.
R-judge: Benchmarking safety risk awareness for llm agents
T. Yuan, Z. He, L. Dong, Y. Wang, R. Zhao, T. Xia, L. Xu, B. Zhou, F. Li, Z. Zhang, R. Wang, and G. Liu
Cited in the paper.
X. Wang, Y. Chen, L. Yuan, Y. Zhang, Y. Li, H. Peng, and H. Ji · 2024
Closest in time.
Swe-agent: Agent-computer interfaces enable automated software engineering, 2024
J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press · 2024
Closest in time.
H. Zhang, J. Huang, K. Mei, Y. Yao, Z. Wang, C. Zhan, H. Wang, and Y. Zhang · 2024
Closest in time.
Opencodeinterpreter: Integrating code generation with execution and refinement
T. Zheng, G. Zhang, T. Shen, X. Liu, B. Y. Lin, J. Fu, W. Chen, and X. Yue · 2024
Closest in time.