Fetching the paper…
Reading the bibliography…
While recent code-specific large language models (LLMs) have greatly enhanced their code generation capabilities, the safety of these models remains under-explored, posing potential risks as insecure code generated by these models may introduce vulnerabilities into real-world systems.
A new measure of rank correlation
Kendall, M. G · 1938
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swayamdipta, S., Schwartz, R., Lourie, N., Wang, Y., Hajishirzi, H., Smith, N. A., and Choi, Y · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
An empirical cybersecurity evaluation of github copilot’s code contributions
Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., and Karri, R · 2021
Earlier work this paper cites.
Multi-lingual evaluation of code generation models
Athiwaratkun, B., Gouda, S. K., Wang, Z., Li, X., Tian, Y., Tan, M., Ahmad, W. U., Wang, S., Sun, Q., Shang, M., et al · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Earlier work this paper cites.
Static analysis for aws best practices in python code
Mukherjee, R., Tripp, O., Liblit, B., and Wilson, M · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., and Karri, R · 2022
Earlier work this paper cites.
Securityeval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques
Siddiq, M. L. and Santos, J. C · 2022
Earlier work this paper cites.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Naik, A., Ashok, A., Dhanasekaran, A. S., Arunkumar, A., Stap, D., et al · 2022
Earlier work this paper cites.
Purple llama cyberseceval: A secure coding benchmark for language models
Bhatt, M., Chennabasappa, S., Nikolaidis, C., Wan, S., Evtimov, I., Gabi, D., Song, D., Ahmad, F., Aschermann, C., Fontana, L., et al · 2023
Earlier work this paper cites.
Large language models for code: Security hardening and adversarial testing
He, J. and Vechev, M · 2023
Earlier work this paper cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Lee, H., Phatale, S., Mansoor, H., Lu, K. R., Mesnard, T., Ferret, J., Bishop, C., Hall, E., Carbune, V., and Rastogi, A · 2023
Earlier work this paper cites.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D · 2023
Earlier work this paper cites.
Cwe: common weakness enumerations, 2023
MITRE · 2023
Earlier work this paper cites.
Llmseceval: A dataset of natural language prompts for security evaluations
Tony, C., Mutas, M., Ferreyra, N. E. D., and Scandariato, R · 2023
Earlier work this paper cites.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2023
Cited alongside, same era.
Magicoder: Source code is all you need
Wei, Y., Wang, Z., Liu, J., Ding, Y., and Zhang, L · 2023
Cited alongside, same era.
Data selection for language models via importance resampling
Xie, S. M., Santurkar, S., Ma, T., and Liang, P. S · 2023
Cited alongside, same era.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M., Aneja, J., Awadalla, H., Awadallah, A., Awan, A. A., Bach, N., Bahree, A., Bakhtiari, A., Bao, J., Behl, H., et al · 2024
A unified debugging approach via llm-based multi-agent synergy
Lee, C., Xia, C. S., Yang, L., Huang, J.-t., Zhu, Z., Zhang, L., and Lyu, M. R · 2024
Closest in time.
Starcoder 2 and the stack v2: The next generation, 2024
Lozhkov, A., Li, R., Allal, L. B., Cassano, F., Lamy-Poirier, J., Tazi, N., Tang, A., Pykhtar, D., Liu, J., Wei, Y., Liu, T., Tian, M., Kocetkov, D., Zucker, A., Belkada, Y., Wang, Z., Liu, Q., Abulkhanov, D., Paul, I., Li, Z., Li, W.-D., Risdal, M., Li, J., Zhu, J., Zhuo, T. Y., Zheltonozhskii, E., Dade, N. O. O., Yu, W., Krauß, L., Jain, N., Su, Y., He, X., Dey, M., Abati, E., Chai, Y., Muennighoff, N., Tang, X., Oblokulov, M., Akiki, C., Marone, M., Mou, C., Mishra, M., Gu, A., Hui, B., Dao, T., Zebaze, A., Dehaene, O., Patry, N., Xu, C., McAuley, J., Hu, H., Scholak, T., Paquet, S., Robinson, J., Anderson, C. J., Chapados, N., Patwary, M., Tajbakhsh, N., Jernite, Y., Ferrandis, C. M., Zhang, L., Hughes, S., Wolf, T., Guha, A., von Werra, L., and de Vries, H · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Nemotron-4 340b technical report
Adler, B., Agarwal, N., Aithal, A., Anh, D. H., Bhattacharya, P., Brundyn, A., Casper, J., Catanzaro, B., Clay, S., Cohen, J., et al · 2024
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Guo, Z. D., Piot, B., Munos, R., Rowland, M., Valko, M., and Calandriello, D · 2024
Cited alongside, same era.
Infinity instruct, 2024
BAAI · 2024
Cited alongside, same era.
Bai, W., Xuan, K., Huang, P., Wu, Q., Wen, J., Wu, J., and Lu, K · 2024
Cited alongside, same era.
Semgrep*: Improving the limited performance of static application security testing (sast) tools
Bennett, G., Hall, T., Winter, E., and Counsell, S · 2024
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Closest in time.
Code llama: Open foundation models for code, 2024
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Sauvestre, R., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al · 2024
Closest in time.
Source code foundation models are transferable binary analysis knowledge bases
Su, Z., Xu, X., Huang, Z., Zhang, K., and Zhang, X · 2024
Closest in time.
Sanitizing large language models in bug detection with data-flow
Wang, C., Zhang, W., Su, Z., Xu, X., and Zhang, X · 2024
Closest in time.
Selfcodealign: Self-alignment for code generation
Wei, Y., Cassano, F., Liu, J., Ding, Y., Jain, N., Mueller, Z., de Vries, H., Von Werra, L., Guha, A., and Zhang, L · 2024
Closest in time.
Qurating: Selecting high-quality data for training language models
Wettig, A., Gupta, A., Malik, S., and Chen, D · 2024
Closest in time.
Libalchemy: A two-layer persistent summary design for taming third-party libraries in static bug-finding systems
Wu, R., He, Y., Huang, J., Wang, C., Tang, W., Shi, Q., Xiao, X., and Zhang, C · 2024
Closest in time.
Opencodeinterpreter: Integrating code generation with execution and refinement
Zheng, T., Zhang, G., Shen, T., Liu, X., Lin, B. Y., Fu, J., Chen, W., and Yue, X · 2024
Closest in time.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2024
Closest in time.
Claude code, 2025
Anthropic · 2025
Closest in time.
Augment code, 2025
AugmentCode · 2025
Closest in time.
Infer, 2025
Meta · 2025
Closest in time.
Openai codex, 2025
OpenAI · 2025
Closest in time.
Weggli, 2025
Weggli · 2025
Closest in time.
Unleashing the power of generative model in recovering variable names from stripped binary
Xu, X., Zhang, Z., Su, Z., Huang, Z., Feng, S., Ye, Y., Jiang, N., Xie, D., Cheng, S., Tan, L., et al · 2025
Closest in time.
Validating network protocol parsers with traceable rfc document interpretation
Zheng, M., Xie, D., Shi, Q., Wang, C., and Zhang, X · 2025
Closest in time.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence
Zhu, Q., Guo, D., Shao, Z., Yang, D., Wang, P., Xu, R., Wu, Y., Li, Y., Gao, H., Ma, S., et al · 2025
Closest in time.