Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are widely utilized in software engineering (SE) tasks, such as code generation and automated program repair.
The measurement of observer agreement for categorical data
Landis, J · 1977
Earlier work this paper cites.
Identifying and filtering near-duplicate documents
Broder, A. Z · 2000
Earlier work this paper cites.
Defects4j: a database of existing faults to enable controlled testing studies for java programs
Just, R., Jalali, D., and Ernst, M. D · 2014
Earlier work this paper cites.
Towards a big data curated benchmark of inter-project code clones
Svajlenko, J., Islam, J. F., Keivanloo, I., Roy, C. K., and Mia, M. M · 2014
Earlier work this paper cites.
Quixbugs: a multi-lingual program repair benchmark set based on the quixey challenge
Lin, D., Koppel, J., Chen, A., and Solar-Lezama, A · 2017
Earlier work this paper cites.
Mapping language to code in programmatic context
Iyer, S., Konstas, I., Cheung, A., and Zettlemoyer, L · 2018
Earlier work this paper cites.
Bugs.jar: a large-scale, diverse dataset of real-world java bugs
Saha, R. K., Lyu, Y., Lam, W., Yoshida, H., and Prasad, M. R · 2018
Earlier work this paper cites.
Learning to mine aligned code and natural language pairs from stack overflow
Yin, P., Deng, B., Chen, E., Vasilescu, B., and Neubig, G · 2018
Earlier work this paper cites.
Re-factoring based program repair applied to programming assignments
Hu, Y., Ahmed, U. Z., Mechtaev, S., Leong, B., and Roychoudhury, A · 2019
Earlier work this paper cites.
Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks
Zhou, Y., Liu, S., Siow, J. K., Du, X., and Liu, Y · 2019
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., Shou, L., Qin, B., Liu, T., Jiang, D., et al · 2020
Earlier work this paper cites.
Graphcodebert: Pre-training code representations with data flow
Guo, D., Ren, S., Lu, S., Feng, Z., Tang, D., Liu, S., Zhou, L., Duan, N., Svyatkovskiy, A., Fu, S., et al · 2020
Earlier work this paper cites.
Intellicode compose: Code generation using transformer
Svyatkovskiy, A., Deng, S. K., Fu, S., and Sundaresan, N · 2020
Earlier work this paper cites.
Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies
Widyasari, R., Sim, S. Q., Lok, C., Qi, H., Phan, J., Tay, Q., Tan, C., Wee, F., Tan, J. E., Yieh, Y., Goh, B., Thung, F., Kang, H. J., Hoang, T., Lo, D., and Ouh, E. L · 2020
Earlier work this paper cites.
Unified pre-training for program understanding and generation
Ahmad, W. U., Chakraborty, S., Ray, B., and Chang, K.-W · 2021
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M. I., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C. J., Terry, M., Le, Q. V., and Sutton, C · 2021
Earlier work this paper cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, Mar. 2021
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., et al · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Lu, S., Guo, D., Ren, S., Huang, J., Svyatkovskiy, A., Blanco, A., Clement, C., Drain, D., Jiang, D., Tang, D., et al · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B., and Komatsuzaki, A · 2021
Earlier work this paper cites.
Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation
Wang, Y., Wang, W., Joty, S., and Hoi, S. C · 2021
Earlier work this paper cites.
D2A: A dataset built for ai-based vulnerability detection methods using differential analysis
Zheng, Y., Pujar, S., Lewis, B. L., Buratti, L., Epstein, E. A., Yang, B., Laredo, J., Morari, A., and Su, Z · 2021
Earlier work this paper cites.
Gpt-neox-20b: An open-source autoregressive language model
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., et al · 2022
Earlier work this paper cites.
Vul4j: A dataset of reproducible java vulnerabilities geared towards the study of program repair techniques
Bui, Q., Scandariato, R., and Ferreyra, N. E. D · 2022
Earlier work this paper cites.
Multipl-e: A scalable and extensible approach to benchmarking neural code generation
Cassano, F., Gouwar, J., Nguyen, D., Nguyen, S., Phipps-Costin, L., Pinckney, D., Yee, M.-H., Zi, Y., Anderson, C. J., Feldman, M. Q., et al · 2022
Earlier work this paper cites.
Pangu-coder: Program synthesis with function-level language modeling
Christopoulou, F., Lampouras, G., Gritta, M., Zhang, G., Guo, Y., Li, Z., Zhang, Q., Xiao, M., Shen, B., Li, L., et al · 2022
Earlier work this paper cites.
Incoder: A generative model for code infilling and synthesis
Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., Yih, W.-t., Zettlemoyer, L., and Lewis, M · 2022
Earlier work this paper cites.
Aixbench: A code generation benchmark dataset
Hao, Y., Li, G., Liu, Y., Miao, X., Zong, H., Jiang, S., Liu, Y., and He, W · 2022
Earlier work this paper cites.
The stack: 3 tb of permissively licensed source code, 2022
Kocetkov, D., Li, R., Allal, L. B., Li, J., Mou, C., Ferrandis, C. M., Jernite, Y., Mitchell, M., Hughes, S., Wolf, T., Bahdanau, D., von Werra, L., and de Vries, H · 2022
Earlier work this paper cites.
Codereviewer: Pre-training for automating code review activities
Li, Z., Lu, S., Guo, D., Duan, N., Jannu, S., Jenks, G., Majumder, D., Green, J., Svyatkovskiy, A., Fu, S., and Sundaresan, N · 2022
Earlier work this paper cites.
Codegen: An open large language model for code with multi-turn program synthesis
Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C · 2022
Earlier work this paper cites.
Chatgpt: Optimizing language models for dialogue
OpenAI · 2022
Cited alongside, same era.
Securityeval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques
Siddiq, M. L., and Santos, J. C · 2022
Cited alongside, same era.
Natural language processing with transformers
Tunstall, L., Von Werra, L., and Wolf, T · 2022
Cited alongside, same era.
Natural language processing with transformers
Tunstall, L., Von Werra, L., and Wolf, T · 2022
Cited alongside, same era.
A systematic evaluation of large language models of code
Xu, F. F., Alon, U., Neubig, G., and Hellendoorn, V. J · 2022
Cited alongside, same era.
Cert: continual pre-training on sketches for library-oriented code generation
Zan, D., Chen, B., Yang, D., Lin, Z., Kim, M., Guan, B., Wang, Y., Chen, W., and Lou, J.-G · 2022
Yan, W., Liu, H., Wang, Y., Li, Y., Chen, Q., Wang, W., Lin, T., Zhao, W., Zhu, L., Sundaram, H., et al · 2023
Later among the works it cites.
Evaluating pre-trained language models for repairing API misuses
Zhang, T., Irsan, I. C., Thung, F., Lo, D., Sharma, A., and Jiang, L · 2023
Later among the works it cites.
Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x
Zheng, Q., Xia, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Wang, Z., Shen, L., Wang, A., Li, Y., et al · 2023
Later among the works it cites.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic · 2024
Later among the works it cites.
Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Santacoder: don’t reach for the stars!
Allal, L. B., Li, R., Kocetkov, D., Mou, C., Akiki, C., Ferrandis, C. M., Muennighoff, N., Mishra, M., Gu, A., Dey, M., et al · 2023
Cited alongside, same era.
Santacoder: don’t reach for the stars!
Allal, L. B., Li, R., Kocetkov, D., Mou, C., Akiki, C., Ferrandis, C. M., Muennighoff, N., Mishra, M., Gu, A., Dey, M., et al · 2023
Cited alongside, same era.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al · 2023
Cited alongside, same era.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., Hui, B., Ji, L., Li, M., Lin, J., Lin, R., Liu, D., Liu, G., Lu, C., Lu, K., Ma, J., Men, R., Ren, X., Ren, X., Tan, C., Tan, S., Tu, J., Wang, P., Wang, S., Wang, W., Wu, S., Xu, B., Xu, J., Yang, A., Yang, H., Yang, J., Yang, S., Yao, Y., Yu, B., Yuan, H., Yuan, Z., Zhang, J., Zhang, X., Zhang, Y., Zhang, Z., Zhou, C., Zhou, J., Zhou, X., and Zhu, T · 2023
Cited alongside, same era.
Can it edit? evaluating the ability of large language models to follow code editing instructions
Cassano, F., Li, L., Sethi, A., Shinn, N., Brennan-Jones, A., Lozhkov, A., Anderson, C. J., and Guha, A · 2023
Cited alongside, same era.
Balloccu, S., Schmidtová, P., Lango, M., and Dušek, O · 2024
Later among the works it cites.
Mercury: A code efficiency benchmark for code large language models
Du, M., Luu, A. T., Ji, B., Liu, Q., and Ng, S.-K · 2024
Later among the works it cites.
Deepseek-coder: When the large language model meets programming – the rise of code intelligence, 2024
et al., D. G · 2024
Later among the works it cites.
IJaDataset 2.0 from Ambient Software Evolution Group
Group, A. S. E · 2024
Later among the works it cites.
Codeeditorbench: Evaluating code editing capability of large language models
Guo, J., Li, Z., Liu, X., Ma, K., Zheng, T., Yu, Z., Pan, D., Li, Y., Liu, R., Wang, Y., et al · 2024
Later among the works it cites.
Exploring the potential of chatgpt in automated code refinement: An empirical study
Guo, Q., Cao, J., Xie, X., Liu, S., Li, X., Chen, B., and Peng, X · 2024
Later among the works it cites.
Large language models for software engineering: A systematic literature review
Hou, X., Zhao, Y., Liu, Y., Yang, Z., Wang, K., Li, L., Luo, X., Lo, D., Grundy, J., and Wang, H · 2024
Later among the works it cites.
Large language models for software engineering: A systematic literature review, 2024
Hou, X., Zhao, Y., Liu, Y., Yang, Z., Wang, K., Li, L., Luo, X., Lo, D., Grundy, J., and Wang, H · 2024
Later among the works it cites.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Jain, N., Han, K., Gu, A., Li, W., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I · 2024
Later among the works it cites.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R · 2024
Later among the works it cites.
Evocodebench: An evolving code generation benchmark aligned with real-world code repositories
Li, J., Li, G., Zhang, X., Dong, Y., and Jin, Z · 2024
Later among the works it cites.
Exploring the effectiveness of llms in automated logging statement generation: An empirical study
Li, Y., Huo, Y., Jiang, Z., Zhong, R., He, P., Su, Y., Briand, L. C., and Lyu, M. R · 2024
Later among the works it cites.
Refining chatgpt-generated code: Characterizing and mitigating code quality issues
Liu, Y., Le-Cong, T., Widyasari, R., Tantithamthavorn, C., Li, L., Le, X. D., and Lo, D · 2024
Later among the works it cites.
Starcoder 2 and the stack v2: The next generation
Lozhkov, A., Li, R., Allal, L. B., Cassano, F., Lamy-Poirier, J., Tazi, N., Tang, A., Pykhtar, D., Liu, J., Wei, Y., et al · 2024
Later among the works it cites.
On inter-dataset code duplication and data leakage in large language models, 2024
López, J. A. H., Chen, B., Saaz, M., Sharma, T., and Varró, D · 2024
Later among the works it cites.
On leakage of code generation evaluation datasets, 2024
Matton, A., Sherborne, T., Aumiller, D., Tommasone, E., Alizadeh, M., He, J., Ma, R., Voisin, M., Gilsenan-McMahon, E., and Gallé, M · 2024
Later among the works it cites.
Introducing Meta Llama 3: The most capable openly available LLM to date
Meta · 2024
Later among the works it cites.
Can large language models write parallel code?
Nichols, D., Davis, J. H., Xie, Z., Rajaram, A., and Bhatele, A · 2024
Later among the works it cites.
Quantifying contamination in evaluating code generation capabilities of language models, 2024
Riddell, M., Ni, A., and Cohan, A · 2024
Later among the works it cites.
Uncovering the limits of machine learning for automatic vulnerability detection
Risse, N., and Böhme, M · 2024
Later among the works it cites.
Gitbug-java: A reproducible benchmark of recent java bugs
Silva, A., Saavedra, N., and Monperrus, M · 2024
Later among the works it cites.
Debugbench: Evaluating debugging capability of large language models
Tian, R., Ye, Y., Qin, Y., Cong, X., Lin, Y., Pan, Y., Wu, Y., Hui, H., Liu, W., Liu, Z., and Sun, M · 2024
Later among the works it cites.
Xie, R., Zeng, Z., Yu, Z., Gao, C., Zhang, S., and Ye, W · 2024
Later among the works it cites.
Codebenchgen: Creating scalable execution-based code generation benchmarks
Xie, Y., Xie, A., Sheth, D., Liu, P., Fried, D., and Rose, C · 2024
Later among the works it cites.
Pythonsaga: Redefining the benchmark to evaluate code generating llms
Yadav, A., Beniwal, H., and Singh, M · 2024
Later among the works it cites.
Unveiling memorization in code models
Yang, Z., Zhao, Z., Wang, C., Shi, J., Kim, D., Han, D., and Lo, D · 2024
Later among the works it cites.
Codereval: A benchmark of pragmatic code generation with generative pre-trained models
Yu, H., Shen, B., Ran, D., Zhang, J., Zhang, Q., Ma, Y., Liang, G., Li, Y., Wang, Q., and Xie, T · 2024
Later among the works it cites.
Out of sight, out of mind: Better automatic vulnerability repair by broadening input ranges and sources
Zhou, X., Kim, K., Xu, B., Han, D., and Lo, D · 2024
Later among the works it cites.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Zhuo, T. Y., Vu, M. C., Chim, J., Hu, H., Yu, W., Widyasari, R., Yusuf, I. N. B., Zhan, H., He, J., Paul, I., et al · 2024
Later among the works it cites.