Fetching the paper…
Reading the bibliography…
Automatic program generation has long been a fundamental challenge in computer science.
Abstract interpretation: A unified lattice model for static analysis of programs by construction or approximation of fixpoints
Cousot, P. and Cousot, R · 1977
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Measuring coding challenge competence with APPS
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Securityeval dataset: Mining vulnerability examples to evaluate machine learning-based code generation techniques
Siddiq, M. L. and Santos, J. C. S · 2022
Earlier work this paper cites.
Purple llama cyberseceval: A secure coding benchmark for language models
Bhatt, M., Chennabasappa, S., Nikolaidis, C., Wan, S., Evtimov, I., Gabi, D., Song, D., Ahmad, F., Aschermann, C., Fontana, L., et al · 2023
Earlier work this paper cites.
DS-1000: A natural and reliable benchmark for data science code generation
Lai, Y., Li, C., Wang, Y., Zhang, T., Zhong, R., Zettlemoyer, L., Yih, W., Fried, D., Wang, S. I., and Yu, T · 2023
Earlier work this paper cites.
Execution-based evaluation for open-domain code generation
Wang, Z., Zhou, S., Fried, D., and Neubig, G · 2023
Earlier work this paper cites.
"false negative-that one is going to kill you": Understanding industry perspectives of static analysis based security testing
Ami, A. S., Moran, K., Poshyvanyk, D., and Nadkarni, A · 2024
Earlier work this paper cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Earlier work this paper cites.
Constrained decoding for secure code generation
Fu, Y., Baker, E., and Chen, Y · 2024
Earlier work this paper cites.
Redcode: Risky code execution and generation benchmark for code agents
Guo, C., Liu, X., Xie, C., Zhou, A., Zeng, Y., Lin, Z., Song, D., and Li, B · 2024
Earlier work this paper cites.
Codelmsec benchmark: Systematically evaluating and finding security vulnerabilities in black-box code language models
Hajipour, H., Hassler, K., Holz, T., Schönherr, L., and Fritz, M · 2024
Earlier work this paper cites.
Instruction tuning for secure code generation
He, J., Vero, M., Krasnopolska, G., and Vechev, M. T · 2024
Earlier work this paper cites.
Qwen2.5-coder technical report
Hui, B., Yang, J., Cui, Z., Yang, J., Liu, D., Zhang, L., Liu, T., Zhang, J., Yu, B., Lu, K., et al · 2024
Earlier work this paper cites.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Cited alongside, same era.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al · 2024
Cited alongside, same era.
Practical attacks against black-box code completion engines, 2024
Jenko, S., He, J., Mündler, N., Vero, M., and Vechev, M · 2024
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R · 2024
Cited alongside, same era.
Automatic programming: Large language models and beyond
Lyu, M. R., Ray, B., Roychoudhury, A., Tan, S. H., and Thongtanunam, P · 2024
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Zhuo, T. Y., Vu, M. C., Chim, J., Hu, H., Yu, W., Widyasari, R., Yusuf, I. N. B., Zhan, H., He, J., Paul, I., et al · 2024
Later among the works it cites.
Claude 3.5 sonnet
Anthropic · 2025
Closest in time.
Model card claude 3 addendum
Anthropic · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Swe-lancer: Can frontier llms earn $1 million from real-world freelance software engineering?
Miserendino, S., Wang, M., Patwardhan, T., and Heidecke, J · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Octopack: Instruction tuning code large language models
Muennighoff, N., Liu, Q., Zebaze, A. R., Zheng, Q., Hui, B., Zhuo, T. Y., Singh, S., Tang, X., von Werra, L., and Longpre, S · 2024
Cited alongside, same era.
Swt-bench: Testing and validating real-world bug-fixes with code agents
Mündler, N., Müller, M. N., He, J., and Vechev, M. T · 2024
Cited alongside, same era.
NYU CTF dataset: A scalable open-source benchmark dataset for evaluating llms in offensive security
Shao, M., Jancheska, S., Udeshi, M., Dolan-Gavitt, B., Xi, H., Milner, K., Chen, B., Yin, M., Garg, S., Krishnamurthy, P., Khorrami, F., Karri, R., and Shafique, M · 2024
Cited alongside, same era.
Barriers to using static application security testing (SAST) tools: A literature review
Wadhams, Z. D., Izurieta, C., and Reinhold, A. M · 2024
Cited alongside, same era.
Openhands: An open platform for ai software developers as generalist agents
Wang, X., Li, B., Song, Y., Xu, F. F., Tang, X., Zhuge, M., Pan, J., Song, Y., Li, B., Singh, J., et al · 2024
Cited alongside, same era.
Prosec: Fortifying code llms with proactive security alignment
Xu, X., Su, Z., Guo, J., Zhang, K., Wang, Z., and Zhang, X · 2024
Cited alongside, same era.
Cybench: A framework for evaluating cybersecurity capabilities and risks of language models
Zhang, A. K., Perry, N., Dulepet, R., Ji, J., Menders, C., Lin, J. W., Jones, E., Hussein, G., Liu, S., Jasper, D., et al · 2024
Cited alongside, same era.
Codestral: Hello, world!
Mistral AI · 2025
Closest in time.
2024 CWE top 25 most dangerous software weaknesses, 2024
MITRE · 2025
Closest in time.
Openai o3-mini system card
OpenAI · 2025
Closest in time.
The openapi specification
OpenAPI Initiative · 2025
Closest in time.
Owasp top ten, 2025
OWASP · 2025
Closest in time.
Cweval: Outcome-driven evaluation on functionality and security of llm code generation
Peng, J., Cui, L., Huang, K., Yang, J., and Ray, B · 2025
Closest in time.
Snyk code: Developer-focused, real-time sast
Snyk · 2025
Closest in time.
2024 developer survey
StackOverflow · 2025
Closest in time.