Fetching the paper…
Reading the bibliography…
With the increasing utilization of large language models such as ChatGPT during software development, it has become crucial to verify the quality of code content it generates.
R. Just, D. Jalali, and M. D. Ernst, “Defects4j: A database of existing faults to enable controlled testing studies for java programs,” in Proceedings of the 23rd International Symposium on Software Testing and Analysis , 2014, pp. 437–440
2014
Earlier work this paper cites.
C. Le Goues, N. Holtschulte, E. K. Smith, Y. Brun, P. Devanbu, S. Forrest, and W. Weimer, “The manybugs and introclass benchmarks for automated repair of c programs,” IEEE Transactions on Software Engineering , vol. 41, pp. 1236–1256, 2015
2015
Earlier work this paper cites.
Q. Xin, “Towards addressing the patch overfitting problem,” in Proceedings of the 39th IEEE/ACM International Conference on Software Engineering , 2017, pp. 489–490
2017
Earlier work this paper cites.
D. Lin, J. Koppel, A. Chen, and A. Solar-Lezama, “Quixbugs: A multi-lingual program repair benchmark set based on the quixey challenge,” in Proceedings Companion of the ACM International Conference on Systems, Programming, Languages, and Applications: Software for Humanity , 2017, pp. 55–56
2017
Earlier work this paper cites.
Y. Tian, K. Pei, S. Jana, and B. Ray, “Deeptest: Automated testing of deep-neural-network-driven autonomous cars,” in Proceedings of the 40th International Conference on Software Engineering , 2018, pp. 303–314
2018
Earlier work this paper cites.
Z. Xu, S. Pang, T. Zhang, X.-P. Luomanu, J. Liu, Y.-T. Tang, X. Yu, and L. Xue, “Cross project defect prediction via balanced distribution adaptation based transfer learning,” Journal of Computer Science and Technology , vol. 34, pp. 1039–1062, 2019
2019
Earlier work this paper cites.
S. Gupta, P. He, C. Meister, and Z. Su, “Machine translation testing via pathological invariance,” in Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2020, pp. 863–875
2020
Earlier work this paper cites.
P. He, C. Meister, and Z. Su, “Structure-invariant testing for machine translation,” in Proceedings of the 42nd ACM/IEEE International Conference on Software Engineering , 2020, pp. 961–973
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
T. M. C. (MITRE), “2021 cwe top 25 most dangerous software weaknesses,” https://cwe.mitre.org/data/definitions/1337.html , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
B. Wang, C. Xu, S. Wang, Z. Gan, Y. Cheng, J. Gao, A. H. Awadallah, and B. Li, “Adversarial GLUE: A multi-task benchmark for robustness evaluation of language models,” in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 , 2021
2021
Earlier work this paper cites.
S. Chen, S. Jin, and X. Xie, “Testing your question answering software via asking recursively,” in Proceedings of the 36th IEEE/ACM International Conference on Automated Software Engineering , 2021, pp. 104–116
2021
Earlier work this paper cites.
S. Feng, J. Keung, X. Yu, Y. Xiao, K. E. Bennin, M. A. Kabir, and M. Zhang, “Coste: Complexity-based oversampling technique to alleviate the class imbalance problem in software defect prediction,” Information and Software Technology , vol. 129, p. 106432, 2021
2021
Earlier work this paper cites.
P. Vaithilingam, T. Zhang, and E. L. Glassman, “Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,” in Chi conference on human factors in computing systems extended abstracts , 2022, pp. 1–7
2022
Earlier work this paper cites.
H. Pearce, B. Ahmad, B. Tan, B. Dolan-Gavitt, and R. Karri, “Asleep at the keyboard? assessing the security of github copilot’s code contributions,” in Proceedings of the 43rd IEEE Symposium on Security and Privacy , 2022, pp. 754–768
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Q. Shen, J. Chen, J. M. Zhang, H. Wang, S. Liu, and M. Tian, “Natural test generation for precise testing of question answering software,” in Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering , 2022, pp. 1–12
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Feng, S. Vanam, M. Cherukupally, W. Zheng, M. Qiu, and H. Chen, “Investigating code generation performance of chat-gpt with crowdsourcing social data,” in Proceedings of the 47th IEEE Computer Software and Applications Conference , 2023, pp. 876–885
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
F. Zhang, B. Chen, Y. Zhang, J. Liu, D. Zan, Y. Mao, J.-G. Lou, and W. Chen, “Repocoder: Repository-level code completion through iterative retrieval and generation,” in Proceedings of the 28th Conference on Empirical Methods in Natural Language Processing , 2023, pp. 2471–2484
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
D. Sobania, M. Briesch, C. Hanna, and J. Petke, “An analysis of the automatic bug fixing performance of chatgpt,” in Proceedings of the IEEE/ACM International Workshop on Automated Program Repair , 2023, pp. 23–30
2023
Cited alongside, same era.
F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman et al. , “Multipl-e: A scalable and polyglot approach to benchmarking neural code generation,” IEEE Transactions on Software Engineering , vol. 49, pp. 3675–3691, 2023
2023
Cited alongside, same era.
B. Athiwaratkun, S. K. Gouda, Z. Wang, X. Li, Y. Tian, M. Tan, W. U. Ahmad, S. Wang, Q. Sun, M. Shang et al. , “Multi-lingual evaluation of code generation models,” in Proceedings of the Eleventh International Conference on Learning Representations , 2023
2023
Cited alongside, same era.
Q. Zheng, X. Xia, X. Zou, Y. Dong, S. Wang, Y. Xue, Z. Wang, L. Shen, A. Wang, Y. Li et al. , “Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x,” in Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2023, pp. 5673–5684
2023
Later among the works it cites.
2023
Later among the works it cites.
E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn, “Detectgpt: Zero-shot machine-generated text detection using probability curvature,” in International Conference on Machine Learning , vol. 202, 2023, pp. 24 950–24 962
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
N. Jiang, K. Liu, T. Lutellier, and L. Tan, “Impact of code language models on automated program repair,” in Proceedings of the 45th IEEE/ACM International Conference on Software Engineering , 2023, pp. 1430–1442
2023
Cited alongside, same era.
R. Khoury, A. R. Avila, J. Brunelle, and B. M. Camara, “How secure is code generated by chatgpt?” in Proceedings of the 2023 IEEE International Conference on Systems, Man, and Cybernetics , 2023, pp. 2445–2451
2023
Cited alongside, same era.
A. Mastropaolo, L. Pascarella, E. Guglielmi, M. Ciniselli, S. Scalabrino, R. Oliveto, and G. Bavota, “On the robustness of code generation techniques: An empirical study on github copilot,” in Proceedings of the 45th IEEE/ACM International Conference on Software Engineering , 2023, pp. 2149–2160
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
T. Y. Zhuo, Z. Li, Y. Huang, F. Shiri, W. Wang, G. Haffari, and Y.-F. Li, “On robustness of prompt-based semantic parsing with large pre-trained language model: An empirical study on codex,” in Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , 2023, pp. 1090–1102
2023
Cited alongside, same era.
M. Fu, C. K. Tantithamthavorn, V. Nguyen, and T. Le, “Chatgpt for vulnerability detection, classification, and repair: How far are we?” in Proceedings of the 30th Asia-Pacific Software Engineering Conference , 2023, pp. 632–636
2023
Cited alongside, same era.
M. D. Purba, A. Ghosh, B. J. Radford, and B. Chu, “Software vulnerability detection using large language models,” in Proceedings of the 34th IEEE International Symposium on Software Reliability Engineering Workshops , 2023, pp. 112–119
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Mündler, J. He, S. Jenko, and M. Vechev, “Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation,” in Proceedings of the Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
G. Inc., “Codeql documentation,” https://codeql.github.com/docs/ , accessed: April 2024
2024
Closest in time.
X. Zhou, T. Zhang, and D. Lo, “Large language model for vulnerability detection: Emerging results and future directions,” in Proceedings of the 44th ACM/IEEE International Conference on Software Engineering: New Ideas and Emerging Results , 2024, pp. 47–51
2024
Closest in time.
A. Madaan, A. Shypula, U. Alon, M. Hashemi, P. Ranganathan, Y. Yang, G. Neubig, and A. Yazdanbakhsh, “Learning performance-improving code edits,” in Proceedings of the Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
Z. Ali, H. Darwis, L. B. Ilmawan, S. R. Jabir, A. R. Manga et al. , “Memory efficient with parameter efficient fine-tuning for code generation using quantization,” in Proceedings of the 18th International Conference on Ubiquitous Information Management and Communication , 2024, pp. 1–6
2024
Closest in time.
T. Liu, C. Xu, and J. McAuley, “Repobench: Benchmarking repository-level code auto-completion systems,” in Proceedings of the Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
2024
Closest in time.
X. Chen, M. Lin, N. Schärli, and D. Zhou, “Teaching large language models to self-debug,” 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Kang, G. An, and S. Yoo, “A quantitative and qualitative evaluation of llm-based explainable fault localization,” Proceedings of the ACM on Software Engineering , vol. 1, pp. 1424–1446, 2024
2024
Closest in time.
S. Feng and C. Chen, “Prompting is all you need: Automated android bug replay with large language models,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , 2024, pp. 1–13
2024
Closest in time.
2024
Closest in time.
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin et al. , “A survey on large language model based autonomous agents,” Frontiers of Computer Science , vol. 18, no. 6, p. 186345, 2024
2024
Closest in time.