Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated unprecedented capabilities in code generation.
P. Jaccard, “Étude comparative de la distribution florale dans une portion des alpes et des jura,” Bull Soc Vaudoise Sci Nat , vol. 37, pp. 547–579, 1901
1901
Earlier work this paper cites.
V. I. Levenshtein et al. , “Binary codes capable of correcting deletions, insertions, and reversals,” in Soviet physics doklady , vol. 10, no. 8. Soviet Union, 1966, pp. 707–710
1966
Earlier work this paper cites.
J. L. Fleiss, “Measuring nominal scale agreement among many raters.” Psychological bulletin , vol. 76, no. 5, p. 378, 1971
1971
Earlier work this paper cites.
J. R. Landis and G. G. Koch, “The measurement of observer agreement for categorical data,” biometrics , pp. 159–174, 1977
1977
Earlier work this paper cites.
R. Chillarege, W.-L. Kao, and R. G. Condit, “Defect type and its impact on the growth curve.” in ICSE , vol. 91, 1991, pp. 246–255
1991
Earlier work this paper cites.
R. Chillarege, I. S. Bhandari et al. , “Orthogonal defect classification-a concept for in-process measurements,” IEEE Transactions on software Engineering , vol. 18, no. 11, pp. 943–956, 1992
1992
Earlier work this paper cites.
A. L. Strauss and J. Corbin, “Open coding,” Social research methods: A reader , pp. 303–306, 2004
2004
Earlier work this paper cites.
V. Braun and V. Clarke, “Using thematic analysis in psychology,” Qualitative research in psychology , vol. 3, no. 2, pp. 77–101, 2006
2006
Earlier work this paper cites.
R. Abreu, P. Zoeteweij, and A. J. Van Gemund, “On the accuracy of spectrum-based fault localization,” in Testing: Academic and industrial conference practice and research techniques-MUTATION (TAICPART-MUTATION 2007) . IEEE, 2007, pp. 89–98
2007
Earlier work this paper cites.
R. Abreu, P. Zoeteweij, and A. J. Van Gemund, “Spectrum-based multiple fault localization,” in 2009 IEEE/ACM International Conference on Automated Software Engineering . IEEE, 2009, pp. 88–99
2009
Earlier work this paper cites.
K. Pan, S. Kim, and E. J. Whitehead, “Toward an understanding of bug fix patterns,” Empirical Software Engineering , vol. 14, pp. 286–315, 2009
2009
Earlier work this paper cites.
C. Le Goues, N. Holtschulte, E. K. Smith, Y. Brun, P. Devanbu, S. Forrest, and W. Weimer, “The manybugs and introclass benchmarks for automated repair of c programs,” IEEE Transactions on Software Engineering , vol. 41, no. 12, pp. 1236–1256, 2015
2015
Earlier work this paper cites.
S. Mechtaev, J. Yi, and A. Roychoudhury, “Angelix: Scalable multiline program patch synthesis via symbolic analysis,” in Proceedings of the 38th international conference on software engineering , 2016, pp. 691–701
2016
Earlier work this paper cites.
Q. Hanam, F. S. d. M. Brito, and A. Mesbah, “Discovering bug patterns in javascript,” in Proceedings of the 2016 24th ACM SIGSOFT international symposium on foundations of software engineering , 2016, pp. 144–156
2016
Earlier work this paper cites.
P. S. Kochhar, D. Wijedasa, and D. Lo, “A large scale study of multiple programming languages and code quality,” in 2016 IEEE 23rd International Conference on Software Analysis, Evolution, and Reengineering (SANER) , vol. 1. IEEE, 2016, pp. 563–573
2016
Earlier work this paper cites.
J. Lazar, J. H. Feng, and H. Hochheiser, Research methods in human-computer interaction . Morgan Kaufmann, 2017
2017
Earlier work this paper cites.
Y. Xiong, J. Wang, R. Yan, J. Zhang, S. Han, G. Huang, and L. Zhang, “Precise condition synthesis for program repair,” in 2017 IEEE/ACM 39th International Conference on Software Engineering (ICSE) . IEEE, 2017, pp. 416–426
2017
Earlier work this paper cites.
S. H. Tan, J. Yi, S. Mechtaev, A. Roychoudhury et al. , “Codeflaws: a programming competition benchmark for evaluating automated program repair tools,” in 2017 IEEE/ACM 39th International Conference on Software Engineering Companion (ICSE-C) . IEEE, 2017, pp. 180–182
2017
Earlier work this paper cites.
D. Lin et al. , “Quixbugs: A multi-lingual program repair benchmark set based on the quixey challenge,” in Proceedings Companion of the 2017 ACM SIGPLAN international conference on systems, programming, languages, and applications: software for humanity , 2017, pp. 55–56
2017
Earlier work this paper cites.
J. Jiang, Y. Xiong, H. Zhang, Q. Gao, and X. Chen, “Shaping program repair space with existing patches and similar code,” in Proceedings of the 27th ACM SIGSOFT international symposium on software testing and analysis , 2018, pp. 298–309
2018
Earlier work this paper cites.
C. Niu, C. Li et al. , “Spt-code: Sequence-to-sequence pre-training for learning source code representations,” in Proceedings of the 44th international conference on software engineering , 2022, pp. 2006–2018
2018
Earlier work this paper cites.
M. Wen, J. Chen, Y. Tian, R. Wu, D. Hao, S. Han, and S.-C. Cheung, “Historical spectrum based fault localization,” IEEE Transactions on Software Engineering , vol. 47, no. 11, pp. 2348–2368, 2019
2019
Earlier work this paper cites.
S. Saha et al. , “Harnessing evolution for multi-hunk program repair,” in 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE) . IEEE, 2019, pp. 13–24
2019
Earlier work this paper cites.
A. Ghanbari, S. Benton, and L. Zhang, “Practical program repair via bytecode mutation,” in Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis , 2019, pp. 19–30
2019
Earlier work this paper cites.
K. Liu, A. Koyuncu, D. Kim, and T. F. Bissyandé, “Tbar: Revisiting template-based automated program repair,” in Proceedings of the 28th ACM SIGSOFT international symposium on software testing and analysis , 2019, pp. 31–42
2019
Earlier work this paper cites.
X. Li, W. Li, Y. Zhang, and L. Zhang, “Deepfl: Integrating multiple fault diagnosis dimensions for deep fault localization,” in Proceedings of the 28th ACM SIGSOFT international symposium on software testing and analysis , 2019, pp. 169–180
2019
Earlier work this paper cites.
P. Gyimesi, B. Vancsics, A. Stocco et al. , “Bugsjs: a benchmark of javascript bugs,” in 2019 12th IEEE Conference on Software Testing, Validation and Verification (ICST) . IEEE, 2019, pp. 90–101
2019
Earlier work this paper cites.
W. Ahmad, S. Chakraborty, B. Ray, and K.-W. Chang, “A transformer-based approach for source code summarization,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 4998–5007
2020
Earlier work this paper cites.
Z. Feng et al. , “CodeBERT: A pre-trained model for programming and natural languages,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , Online, Nov. 2020, pp. 1536–1547
2020
Earlier work this paper cites.
B. Roziere et al. , “Unsupervised translation of programming languages,” Advances in neural information processing systems , vol. 33, pp. 20 601–20 611, 2020
2020
Earlier work this paper cites.
A. Rahman, E. Farhana, C. Parnin, and L. Williams, “Gang of eight: A defect taxonomy for infrastructure as code scripts,” in Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering , 2020, pp. 752–764
2020
Earlier work this paper cites.
N. Humbatova, G. Jahangirova, G. Bavota, V. Riccio, A. Stocco, and P. Tonella, “Taxonomy of real faults in deep learning systems,” in Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering , 2020, pp. 1110–1121
2020
Cited alongside, same era.
2021
Cited alongside, same era.
A. Antoine, S. Malacria, N. Marquardt, and G. Casiez, “Interaction illustration taxonomy: Classification of styles and techniques for visually representing interaction scenarios,” in Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , 2021, pp. 1–22
2021
Cited alongside, same era.
Y. Li, S. Wang, and T. Nguyen, “Fault localization with code coverage representation learning,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 2021, pp. 661–673
2021
B. Chen, F. Zhang, A. Nguyen, D. Zan, Z. Lin, J.-G. Lou, and W. Chen, “Codet: Code generation with generated tests,” in The Eleventh International Conference on Learning Representations , 2023
2023
Later among the works it cites.
T. X. Olausson, J. P. Inala, C. Wang, J. Gao, and A. Solar-Lezama, “Is self-repair a silver bullet for code generation?” in The Twelfth International Conference on Learning Representations , 2023
2023
Later among the works it cites.
J. Li et al. , “Skcoder: A sketch-based approach for automatic code generation,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 2023, pp. 2124–2135
2023
Later among the works it cites.
S. Arakelyan, R. Das, Y. Mao, and X. Ren, “Exploring distributional shifts in large language models for code analysis,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 16 298–16 314
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2021
Cited alongside, same era.
D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt, “Measuring coding challenge competence with apps,” NeurIPS , 2021
2021
Cited alongside, same era.
E. Nijkamp, B. Pang, H. Hayashi et al. , “Codegen: An open large language model for code with multi-turn program synthesis,” in The Eleventh International Conference on Learning Representations , 2022
2022
Cited alongside, same era.
D. Fried, A. Aghajanyan et al. , “Incoder: A generative model for code infilling and synthesis,” in The Eleventh International Conference on Learning Representations , 2022
2022
Cited alongside, same era.
P. Vaithilingam, T. Zhang, and E. L. Glassman, “Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,” in Chi conference on human factors in computing systems extended abstracts , 2022, pp. 1–7
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Y. Li, S. Wang, and T. N. Nguyen, “Dear: A novel deep learning-based approach for automated program repair,” in Proceedings of the 44th International Conference on Software Engineering , 2022, pp. 511–523
2022
Cited alongside, same era.
H. Ye, M. Martinez, X. Luo, T. Zhang, and M. Monperrus, “Selfapr: Self-supervised program repair with test execution diagnostics,” in Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering , 2022, pp. 1–13
2022
Cited alongside, same era.
K. Jesse, T. Ahmed, P. T. Devanbu, and E. Morgan, “Large language models and simple, stupid bugs,” in IEEE/ACM 20th International Conference on Mining Software Repositories (MSR) , 2023, pp. 563–575
2023
Later among the works it cites.
Y. Liu, T. Le-Cong et al. , “Refining chatgpt-generated code: Characterizing and mitigating code quality issues,” ACM Transactions on Software Engineering and Methodology , 2023
2023
Later among the works it cites.
A. Rahman, D. B. Bose, R. Shakya, and R. Pandita, “Come for syntax, stay for speed, understand defects: an empirical study of defects in julia programs,” Empirical Software Engineering , vol. 28, no. 4, p. 93, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Y. Zhuo, “Ice-score: Instructing large language models to evaluate code,” in Findings of the Association for Computational Linguistics: EACL 2024 , 2024, pp. 2232–2242
2024
Closest in time.
W. Tong and T. Zhang, “Codejudge: Evaluating code generation with large language models,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , 2024, pp. 20 032–20 051
2024
Closest in time.
Q. Zhu, Q. Liang, Z. Sun, Y. Xiong, L. Zhang, and S. Cheng, “Grammart5: Grammar-integrated pretrained encoder-decoder neural model for code,” in Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , 2024, pp. 1–13
2024
Closest in time.
Y. Cai, Y. Lin, C. Liu et al. , “On-the-fly adapting code summarization on trainable cost-effective language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Y. Ding, B. Steenhoek, K. Pei, G. Kaiser, W. Le, and B. Ray, “Traced: Execution-aware pre-training for source code,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , 2024, pp. 1–12
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Luo, C. Xu et al. , “Wizardcoder: Empowering code large language models with evol-instruct,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
R. Bairi, A. Sonwane et al. , “Codeplan: Repository-level coding using llms and planning,” in Proceedings of the 32nd ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2024
2024
Closest in time.
X. Chen, M. Lin, N. Schärli, and D. Zhou, “Teaching large language models to self-debug,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
Y. Ding, M. J. Min, G. Kaiser, and B. Ray, “Cycle: Learning to self-refine the code generation,” Proceedings of the ACM on Programming Languages , vol. 8, no. OOPSLA1, pp. 392–418, 2024
2024
Closest in time.
A. Madaan et al. , “Self-refine: Iterative refinement with self-feedback,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Y. Liu, C. Tantithamthavorn, Y. Liu, and L. Li, “On the reliability and explainability of automated code generation approaches,” ACM Transactions on Software Engineering and Methodology , 2024
2024
Closest in time.
2024
Closest in time.
B. Kou, S. Chen, Z. Wang, L. Ma, and T. Zhang, “Do large language models pay similar attention like human programmers when generating code?” PACMSE , vol. FSE, no. 100, 2024
2024
Closest in time.
Z. Liu, Y. Tang, X. Luo, Y. Zhou, and L. F. Zhang, “No need to lift a finger anymore? assessing the quality of code generation by chatgpt,” IEEE Transactions on Software Engineering , 2024
2024
Closest in time.
R. Pan, A. R. Ibrahimzada et al. , “Lost in translation: A study of bugs introduced by large language models while translating code,” in Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , 2024, pp. 1–13
2024
Closest in time.
Z. Wang, Z. Zhou, D. Song et al. , “GitHub Repository of The Paper: Towards Understanding Error Characteristics of Code Generated by LLMs.” https://github.com/ma-labo/defects4codellm , 2025
2025
Closest in time.