Fetching the paper…
Reading the bibliography…
We present a method for systematically evaluating the correctness and robustness of instruction-tuned large language models (LLMs) for code generation via a new benchmark, Turbulence.
V. I. Levenshtein, “Binary codes capable of correcting deletions, insertions, and reversals,” Soviet physics. Doklady , vol. 10, pp. 707–710, 1965. [Online]. Available: https://api.semanticscholar.org/CorpusID:60827152
1965
Earlier work this paper cites.
W. M. McKeeman, “Differential testing for software,” Digit. Tech. J. , vol. 10, no. 1, pp. 100–107, 1998. [Online]. Available: https://www.hpl.hp.com/hpjournal/dtj/vol10num1/vol10num1art9.pdf
1998
Earlier work this paper cites.
I. Ciupa, B. Meyer, M. Oriol, and A. Pretschner, “Finding faults: Manual testing vs. random+ testing vs. user reports,” in ISSRE . IEEE Computer Society, 2008, pp. 157–166. [Online]. Available: https://doi.org/10.1109/ISSRE.2008.18
2008
Earlier work this paper cites.
NVIDIA, “Determinism in deep learning,” https://developer.download.nvidia.com/video/gputechconf/gtc/2019/presentation/s9911-determinism-in-deep-learning.pdf , dec 2010
2010
Earlier work this paper cites.
T. B. Hashimoto, H. Zhang, and P. Liang, “Unifying human and statistical evaluation for natural language generation,” in NAACL-HLT , J. Burstein, C. Doran, and T. Solorio, Eds. Association for Computational Linguistics, 2019, pp. 1689–1701. [Online]. Available: https://doi.org/10.18653/v1/n19-1169
2019
Earlier work this paper cites.
M. Gardner, Y. Artzi, V. Basmova, J. Berant, B. Bogin, S. Chen, P. Dasigi, D. Dua, Y. Elazar, A. Gottumukkala, N. Gupta, H. Hajishirzi, G. Ilharco, D. Khashabi, K. Lin, J. Liu, N. F. Liu, P. Mulcaire, Q. Ning, S. Singh, N. A. Smith, S. Subramanian, R. Tsarfaty, E. Wallace, A. Zhang, and B. Zhou, “Evaluating models’ local decision boundaries via contrast sets,” in EMNLP , ser. Findings of ACL, T. Cohn, Y. He, and Y. Liu, Eds., vol. EMNLP 2020. Association for Computational Linguistics, 2020, pp. 1307–1323. [Online]. Available: https://doi.org/10.18653/v1/2020.findings-emnlp.117
2020
Earlier work this paper cites.
M. Caccia, L. Caccia, W. Fedus, H. Larochelle, J. Pineau, and L. Charlin, “Language gans falling short,” in ICLR . OpenReview.net, 2020. [Online]. Available: https://openreview.net/forum?id=BJgza6VtPB
2020
Earlier work this paper cites.
H. Zhang, Z. Li, G. Li, L. Ma, Y. Liu, and Z. Jin, “Generating adversarial examples for holding robustness of source code processing models,” in AAAI . AAAI Press, 2020, pp. 1169–1176. [Online]. Available: https://doi.org/10.1609/aaai.v34i01.5469
2020
Earlier work this paper cites.
N. Yefet, U. Alon, and E. Yahav, “Adversarial examples for models of code,” Proc. ACM Program. Lang. , vol. 4, no. OOPSLA, pp. 162:1–162:30, 2020. [Online]. Available: https://doi.org/10.1145/3428230
2020
Earlier work this paper cites.
J. D. Weisz, M. J. Muller, S. Houde, J. T. Richards, S. I. Ross, F. Martinez, M. Agarwal, and K. Talamadupula, “Perfection not required? Human-AI partnerships in code translation,” in IUI . ACM, 2021, pp. 402–412. [Online]. Available: https://doi.org/10.1145/3397481.3450656
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt, “Measuring coding challenge competence with APPS,” in NeurIPS Datasets and Benchmarks , J. Vanschoren and S. Yeung, Eds., 2021. [Online]. Available: https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/c24cd76e1ce41366a4bbe8a49b02a028-Abstract-round2.html
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
L. Reynolds and K. McDonell, “Prompt programming for large language models: Beyond the few-shot paradigm,” in CHI ’21: CHI Conference on Human Factors in Computing Systems , Y. Kitamura, A. Quigley, K. Isbister, and T. Igarashi, Eds. ACM, 2021, pp. 314:1–314:7. [Online]. Available: https://doi.org/10.1145/3411763.3451760
2021
Earlier work this paper cites.
S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. B. Clement, D. Drain, D. Jiang, D. Tang, G. Li, L. Zhou, L. Shou, L. Zhou, M. Tufano, M. Gong, M. Zhou, N. Duan, N. Sundaresan, S. K. Deng, S. Fu, and S. Liu, “CodeXGLUE: A machine learning benchmark dataset for code understanding and generation,” in NeurIPS Datasets and Benchmarks , J. Vanschoren and S. Yeung, Eds., 2021. [Online]. Available: https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/c16a5320fa475530d9583c34fd356ef5-Abstract-round1.html
2021
Earlier work this paper cites.
S. Srikant, S. Liu, T. Mitrovska, S. Chang, Q. Fan, G. Zhang, and U. O’Reilly, “Generating adversarial computer programs using optimized obfuscations,” in ICLR . OpenReview.net, 2021. [Online]. Available: https://openreview.net/forum?id=PH5PH9ZO_4
2021
Earlier work this paper cites.
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” in ICLR . OpenReview.net, 2022. [Online]. Available: https://openreview.net/forum?id=gEZrGCozdqR
2022
Earlier work this paper cites.
N. Nguyen and S. Nadi, “An empirical evaluation of GitHub Copilot’s code suggestions,” in MSR . ACM, 2022, pp. 1–5. [Online]. Available: https://doi.org/10.1145/3524842.3528470
2022
Earlier work this paper cites.
F. F. Xu, U. Alon, G. Neubig, and V. J. Hellendoorn, “A systematic evaluation of large language models of code,” in MAPS@PLDI 2022: 6th ACM SIGPLAN International Symposium on Machine Programming , S. Chaudhuri and C. Sutton, Eds. ACM, 2022, pp. 1–10. [Online]. Available: https://doi.org/10.1145/3520312.3534862
2022
Earlier work this paper cites.
Z. Yang, J. Shi, J. He, and D. Lo, “Natural attack for pre-trained models of code,” in ICSE . ACM, 2022, pp. 1482–1493. [Online]. Available: https://doi.org/10.1145/3510003.3510146
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” in NeurIPS , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., 2022. [Online]. Available: http://papers.nips.cc/paper_files/paper/2022/hash/8bb0d291acd4acf06ef112099c16f326-Abstract-Conference.html
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
H. Zhang, Z. Fu, G. Li, L. Ma, Z. Zhao, H. Yang, Y. Sun, Y. Liu, and Z. Jin, “Towards robustness of deep program processing models - detection, estimation, and enhancement,” ACM Trans. Softw. Eng. Methodol. , vol. 31, no. 3, pp. 50:1–50:40, 2022. [Online]. Available: https://doi.org/10.1145/3511887
2022
Cited alongside, same era.
Z. Zeng, H. Tan, H. Zhang, J. Li, Y. Zhang, and L. Zhang, “An extensive study on pre-trained models for program understanding and generation,” in ISSTA ’22: 31st ACM SIGSOFT , S. Ryu and Y. Smaragdakis, Eds. ACM, 2022, pp. 39–51. [Online]. Available: https://doi.org/10.1145/3533767.3534390
2022
Cited alongside, same era.
M. Wei, Y. Huang, J. Yang, J. Wang, and S. Wang, “CoCoFuzzing: Testing neural code models with coverage-guided fuzzing,” IEEE Trans. Reliab. , vol. 72, no. 3, pp. 1276–1289, 2023. [Online]. Available: https://doi.org/10.1109/TR.2022.3208239
Ollama, “Codellama,” https://ollama.ai/library/codellama , oct 2023
2023
Closest in time.
L. Yuan, Y. Chen, G. Cui, H. Gao, F. Zou, X. Cheng, H. Ji, Z. Liu, and M. Sun, “Revisiting out-of-distribution robustness in NLP: Benchmarks, analysis, and LLMs evaluations,” in NeurIPS , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., 2023. [Online]. Available: http://papers.nips.cc/paper_files/paper/2023/hash/b6b5f50a2001ad1cbccca96e693c4ab4-Abstract-Datasets_and_Benchmarks.html
2023
Closest in time.
OpenAI, “About OpenAI,” https://openai.com/about, 2023
2023
Closest in time.
Meta, “Introducing Code Llama, a state-of-the-art large language model for coding,” https://ai.meta.com/blog/code-llama-large-language-model-coding/ , oct 2023
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
V. Lomshakov, S. V. Kovalchuk, M. Omelchenko, S. I. Nikolenko, and A. Aliev, “Fine-tuning large language models for answering programming questions with code snippets,” in ICCS , ser. Lecture Notes in Computer Science, J. Mikyska, C. de Mulatier, M. Paszynski, V. V. Krzhizhanovskaya, J. J. Dongarra, and P. M. A. Sloot, Eds., vol. 14074. Springer, 2023, pp. 171–179. [Online]. Available: https://doi.org/10.1007/978-3-031-36021-3_15
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Liu, C. S. Xia, Y. Wang, and L. Zhang, “Is your code generated by chatgpt really correct? Rigorous evaluation of large language models for code generation,” in NeurIPS , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., 2023. [Online]. Available: http://papers.nips.cc/paper_files/paper/2023/hash/43e9d647ccd3e4b7b5baab53f0368686-Abstract-Conference.html
2023
Cited alongside, same era.
S. I. Ross, F. Martinez, S. Houde, M. J. Muller, and J. D. Weisz, “The programmer’s assistant: Conversational interaction with a large language model for software development,” in IUI . ACM, 2023, pp. 491–514. [Online]. Available: https://doi.org/10.1145/3581641.3584037
2023
Cited alongside, same era.
N. Perry, M. Srivastava, D. Kumar, and D. Boneh, “Do users write more insecure code with AI assistants?” in Proceedings of the 2023 ACM SIGSAC, CCS , W. Meng, C. D. Jensen, C. Cremers, and E. Kirda, Eds. ACM, 2023, pp. 2785–2799. [Online]. Available: https://doi.org/10.1145/3576915.3623157
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Mastropaolo, L. Pascarella, E. Guglielmi, M. Ciniselli, S. Scalabrino, R. Oliveto, and G. Bavota, “On the robustness of code generation techniques: An empirical study on GitHub Copilot,” in ICSE . IEEE, 2023, pp. 2149–2160. [Online]. Available: https://doi.org/10.1109/ICSE48619.2023.00181
2023
Cited alongside, same era.
Z. Tian, J. Chen, and Z. Jin, “Code difference guided adversarial example generation for deep code models,” in ASE . IEEE, 2023, pp. 850–862. [Online]. Available: https://doi.org/10.1109/ASE56229.2023.00149
2023
Cited alongside, same era.
J. Gawlikowski, C. R. N. Tassi, M. Ali, J. Lee, M. Humt, J. Feng, A. M. Kruspe, R. Triebel, P. Jung, R. Roscher, M. Shahzad, W. Yang, R. Bamler, and X. Zhu, “A survey of uncertainty in deep neural networks,” Artif. Intell. Rev. , vol. 56, no. S1, pp. 1513–1589, 2023. [Online]. Available: https://doi.org/10.1007/s10462-023-10562-9
2023
Closest in time.
Logilab and P. contributors, “Pylint,” https://pylint.pycqa.org/en/latest/index.html#pylint , nov 2023
2023
Closest in time.
S. Thakur, B. Ahmad, Z. Fan, H. Pearce, B. Tan, R. Karri, B. Dolan-Gavitt, and S. Garg, “Benchmarking large language models for automated Verilog RTL code generation,” in DATE . IEEE, 2023, pp. 1–6. [Online]. Available: https://doi.org/10.23919/DATE56975.2023.10137086
2023
Closest in time.
A. M. Dakhel, V. Majdinasab, A. Nikanjam, F. Khomh, M. C. Desmarais, and Z. M. J. Jiang, “GitHub Copilot AI pair programmer: Asset or liability?” J. Syst. Softw. , vol. 203, p. 111734, 2023. [Online]. Available: https://doi.org/10.1016/j.jss.2023.111734
2023
Closest in time.
Z. Li, C. Wang, Z. Liu, H. Wang, D. Chen, S. Wang, and C. Gao, “CCTEST: Testing and repairing code completion systems,” in ICSE . IEEE, 2023, pp. 1238–1250. [Online]. Available: https://doi.org/10.1109/ICSE48619.2023.00110
2023
Closest in time.
J. Jia, S. Srikant, T. Mitrovska, C. Gan, S. Chang, S. Liu, and U. O’Reilly, “ClawSAT: Towards both robust and accurate code models,” in IEEE International Conference on Software Analysis, Evolution and Reengineering, SANER , T. Zhang, X. Xia, and N. Novielli, Eds. IEEE, 2023, pp. 212–223. [Online]. Available: https://doi.org/10.1109/SANER56733.2023.00029
2023
Closest in time.
Z. Tian, J. Chen, and Z. Jin, “Code difference guided adversarial example generation for deep code models,” in ASE . IEEE, 2023, pp. 850–862. [Online]. Available: https://doi.org/10.1109/ASE56229.2023.00149
2023
Closest in time.
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, W. Ye, Y. Zhang, Y. Chang, P. S. Yu, Q. Yang, and X. Xie, “A survey on evaluation of large language models,” ACM Trans. Intell. Syst. Technol. , vol. 15, no. 3, pp. 39:1–39:45, 2024. [Online]. Available: https://doi.org/10.1145/3641289
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Du, M. Liu, K. Wang, H. Wang, J. Liu, Y. Chen, J. Feng, C. Sha, X. Peng, and Y. Lou, “Evaluating large language models in class-level code generation,” in ICSE . ACM, 2024, pp. 81:1–81:13. [Online]. Available: https://doi.org/10.1145/3597503.3639219
2024
Closest in time.
S. S. Rajan, E. O. Soremekun, and S. Chattopadhyay, “Knowledge-based consistency testing of large language models,” in EMNLP . Association for Computational Linguistics, 2024, pp. 10 185–10 196. [Online]. Available: https://aclanthology.org/2024.findings-emnlp.596
2024
Closest in time.
2024
Closest in time.
M. J. Min, Y. Ding, L. Buratti, S. Pujar, G. E. Kaiser, S. Jana, and B. Ray, “Beyond accuracy: Evaluating self-consistency of code large language models with IdentityChain,” in ICLR . OpenReview.net, 2024. [Online]. Available: https://openreview.net/forum?id=caW7LdAALh
2024
Closest in time.
G. Yang, Y. Zhou, W. Yang, T. Yue, X. Chen, and T. Chen, “How important are good method names in neural code generation? A model robustness perspective,” ACM Trans. Softw. Eng. Methodol. , vol. 33, no. 3, pp. 60:1–60:35, 2024. [Online]. Available: https://doi.org/10.1145/3630010
2024
Closest in time.
H. Zhang, S. Lu, Z. Li, Z. Jin, L. Ma, Y. Liu, and G. Li, “CodeBERT-Attack: Adversarial attack against source code deep learning models via pre-trained model,” J. Softw. Evol. Process. , vol. 36, no. 3, 2024. [Online]. Available: https://doi.org/10.1002/smr.2571
2024
Closest in time.
J. Chen, Z. Pan, X. Hu, Z. Li, G. Li, and X. Xia, “Reasoning runtime behavior of a program with LLM: How far are we?” in ICSE 2025 . Los Alamitos, CA, USA: IEEE Computer Society, May 2025, pp. 140–152. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/ICSE55347.2025.00012
2025
Closest in time.
S. Honarvar and A. Donaldson, “ShahinHonarvar/Turbulence-Benchmark: Version 1.0,” Jan. 2025. [Online]. Available: https://doi.org/10.5281/zenodo.14732749
2025
Closest in time.