Fetching the paper…
Reading the bibliography…
Code Large Language Models (Code LLMs) have been increasingly used by developers to boost productivity, but they often generate vulnerable code.
M. Welling and Y. W. Teh, “Bayesian Learning via Stochastic Gradient Langevin dynamics,” in Proceedings of the 28th international conference on machine learning (ICML-11) . Citeseer, 2011, pp. 681–688
2011
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “Guided Open Vocabulary Image Captioning with Constrained Beam Search,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , 2017, pp. 936–945
2017
Earlier work this paper cites.
C. D. V. Hoang, G. Haffari, and T. Cohn, “Towards Decoding as Continuous Optimisation in Neural Machine Translation,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , 2017, pp. 146–156
2017
Earlier work this paper cites.
M. Post and D. Vilar, “Fast Lexically Constrained Decoding with Dynamic Beam Allocation for Neural Machine Translation,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , 2018, pp. 1314–1324
2018
Earlier work this paper cites.
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The Curious Case of Neural Text Degeneration,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
GitHub, “Github Copilot: Your AI Pair Programmer,” https://github.com/features/copilot/ , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 4582–4597
2021
Earlier work this paper cites.
X. Lu, P. West, R. Zellers, R. Le Bras, C. Bhagavatula, and Y. Choi, “NeuroLogic Decoding:(Un) supervised Neural Text Generation with Predicate Logic Constraints,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 4288–4299
2021
Earlier work this paper cites.
S. Kumar, E. Malmi, A. Severyn, and Y. Tsvetkov, “Controlled Text Generation as Continuous Optimization with Multiple Constraints,” in Advances in Neural Information Processing Systems , 2021
2021
Earlier work this paper cites.
Eirini Kalliamvakou, GitHub Blog, “Research: quantifying GitHub Copilot’s impact on developer productivity and happiness,” https://github.blog/2022-09-07-research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness/ , 2022
2022
Earlier work this paper cites.
Maxim Tabachnyk and Stoyan Nikolov, Google Research, “ML-Enhanced Code Completion Improves Developer Productivity,” https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/ , 2022
2022
Earlier work this paper cites.
H. Pearce, B. Ahmad, B. Tan, B. Dolan-Gavitt, and R. Karri, “Asleep at the keyboard? assessing the security of github copilot’s code contributions,” in 2022 IEEE Symposium on Security and Privacy (SP) . IEEE, 2022, pp. 754–768
2022
Earlier work this paper cites.
M. L. Siddiq and J. C. Santos, “Securityeval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques,” in Proceedings of the 1st International Workshop on Mining Software Repositories Applications for Privacy and Security , 2022, pp. 29–33
2022
Earlier work this paper cites.
S. Kumar, B. Paria, and Y. Tsvetkov, “Gradient-based constrained sampling from language models,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Chan Woo Kim, “Guiding Text Generation with Constrained Beam Search in Transformers,” https://huggingface.co/blog/constrained-beam-search , 2022
2022
Cited alongside, same era.
L. Qin, S. Welleck, D. Khashabi, and Y. Choi, “COLD decoding: Energy-based constrained text generation with langevin dynamics,” in Advances in Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
H. Pearce, B. Tan, B. Ahmad, R. Karri, and B. Dolan-Gavitt, “Examining Zero-Shot Vulnerability Repair with Large Language Models,” in 2023 IEEE Symposium on Security and Privacy (SP) . IEEE, 2023, pp. 2339–2356
2023
Later among the works it cites.
A. Storhaug, J. Li, and T. Hu, “Efficient Avoidance of Vulnerabilities in Auto-completed Smart Contract Code Using Vulnerability-constrained Decoding,” in 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE) . IEEE, 2023, pp. 683–693
2023
Later among the works it cites.
H. Hajipour, K. Hassler, T. Holz, L. Schönherr, and M. Fritz, “CodeLMSec Benchmark: Systematically Evaluating and Finding Security Vulnerabilities in Black-Box Code Language Models,” in 2024 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) . IEEE, 2024, pp. 684–709
2024
Closest in time.
J. Liu, C. S. Xia, Y. Wang, and L. Zhang, “Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation,” Advances in Neural Information Processing Systems , vol. 36, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Liu, Z. Yang, T. Tao, X. Liang, J. Bao, Z. Li, X. He, S. Cui, and Z. Hu, “Don’t take it literally: An edit-invariant sequence loss for text generation,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2022
2022
Cited alongside, same era.
Amazon, “Amazon CodeWhisperer: Your AI-powered productivity tool for the IDE and command line ,” https://aws.amazon.com/codewhisperer/ , 2023
2023
Cited alongside, same era.
Tiernan Ray, ZDNet, “Microsoft has over a million paying Github Copilot users: CEO Nadella,” https://www.zdnet.com/article/microsoft-has-over-a-million-paying-github-copilot-users-ceo-nadella/ , 2023
2023
Cited alongside, same era.
J. He and M. Vechev, “Large language models for code: Security hardening and adversarial testing,” in Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security , 2023, pp. 1865–1879
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al. , “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023
2023
Cited alongside, same era.
R. Khoury, A. R. Avila, J. Brunelle, and B. M. Camara, “How secure is code generated by chatgpt?” in 2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC) . IEEE, 2023, pp. 2445–2451
2023
Cited alongside, same era.
N. Tihanyi, T. Bisztray, R. Jain, M. A. Ferrag, L. C. Cordeiro, and V. Mavroeidis, “The FormAI Dataset: Generative AI in Software Security through the Lens of Formal Verification,” in Proceedings of the 19th International Conference on Predictive Models and Data Analytics in Software Engineering , 2023, pp. 33–43
2023
Cited alongside, same era.
2024
Closest in time.
J. He, M. Vero, G. Krasnopolska, and M. Vechev, “Instruction tuning for secure code generation,” in Proceedings of the International Conference on Machine Learning (ICML) , 2024
2024
Closest in time.
OpenAI, “Gpt-4 technical report,” 2024
2024
Closest in time.
2024
Closest in time.
C. Team, “Codegemma: Open code models based on gemma,” 2024. [Online]. Available: https://goo.gle/codegemma
2024
Closest in time.
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li et al. , “Deepseek-coder: When the large language model meets programming – the rise of code intelligence,” 2024
2024
Closest in time.
B. Rozière, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, R. Sauvestre, T. Remez et al. , “Code llama: Open foundation models for code,” 2024
2024
Closest in time.
Y. Fu, P. Liang, A. Tahir, Z. Li, M. Shahin, and J. Yu, “Security Weaknesses of Copilot Generated Code in GitHub,” in ACM Transactions on Software Engineering and Methodology . ACM, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. T. Islam and P. Najafirad, “Code Security Vulnerability Repair Using Reinforcement Learning with Large Language Models,” in Proceedings of the AAAI Conference on Artificial Intelligence Workshop , 2024
2024
Closest in time.