Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) for code have gained significant attention recently.
C. B. Seaman, “Qualitative methods in empirical studies of software engineering,” IEEE Transactions on software engineering , vol. 25, no. 4, pp. 557–572, 1999
1999
Earlier work this paper cites.
A. N. Oppenheim, Questionnaire design, interviewing and attitude measurement . Bloomsbury Publishing, 2000
2000
Earlier work this paper cites.
A. Cater-Steel, M. Toleman, and T. Rout, “Addressing the challenges of replications of surveys in software engineering research,” in 2005 International Symposium on Empirical Software Engineering, 2005. IEEE, 2005, pp. 10–pp
2005
Earlier work this paper cites.
Anthropic, “Model card and evaluations for claude models,” 2013. [Online]. Available: https://www.anthropic.com/index/introducing-claude
2013
Earlier work this paper cites.
R. Just, D. Jalali, L. Inozemtseva, M. D. Ernst, R. Holmes, and G. Fraser, “Are mutants a valid substitute for real faults in software testing?” in Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering , ser. FSE 2014. New York, NY, USA: Association for Computing Machinery, 2014, p. 654–665. [Online]. Available: https://doi.org/10.1145/2635868.2635929
2014
Earlier work this paper cites.
G. Fraser and A. Arcuri, “A large-scale evaluation of automated unit test generation using evosuite,” ACM Trans. Softw. Eng. Methodol. , vol. 24, no. 2, dec 2014. [Online]. Available: https://doi.org/10.1145/2685612
2014
Earlier work this paper cites.
F. Fischer, K. Böttinger, H. Xiao, C. Stransky, Y. Acar, M. Backes, and S. Fahl, “Stack overflow considered harmful? the impact of copy&paste on android application security,” in 2017 IEEE Symposium on Security and Privacy (SP) , 2017, pp. 121–136
2017
Earlier work this paper cites.
S. H. Tan, J. Yi, S. Mechtaev, A. Roychoudhury et al. , “Codeflaws: a programming competition benchmark for evaluating automated program repair tools,” in 2017 IEEE/ACM 39th International Conference on Software Engineering Companion (ICSE-C) . IEEE, 2017, pp. 180–182
2017
Earlier work this paper cites.
D. Spadini, M. Aniche, and A. Bacchelli, “Pydriller: Python framework for mining software repositories,” in Proceedings of the 2018 26th ACM Joint meeting on european software engineering conference and symposium on the foundations of software engineering , 2018, pp. 908–911
2018
Earlier work this paper cites.
T. Zhang, G. Upadhyaya, A. Reinhardt, H. Rajan, and M. Kim, “Are code examples on an online q&a forum reliable?: A study of api misuse on stack overflow,” in 2018 IEEE/ACM 40th International Conference on Software Engineering (ICSE) , 2018, pp. 886–896
2018
Earlier work this paper cites.
M. Monperrus, “Automatic software repair: A bibliography,” ACM Comput. Surv. , vol. 51, no. 1, jan 2018. [Online]. Available: https://doi.org/10.1145/3105906
2018
Earlier work this paper cites.
T. Zhang, C. Gao, L. Ma, M. Lyu, and M. Kim, “An empirical study of common challenges in developing deep learning applications,” in 2019 IEEE 30th International Symposium on Software Reliability Engineering (ISSRE) , 2019, pp. 104–115
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
N. Humbatova, G. Jahangirova, G. Bavota, V. Riccio, A. Stocco, and P. Tonella, “Taxonomy of real faults in deep learning systems,” in Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering , 2020, pp. 1110–1121
2020
Earlier work this paper cites.
M. Verdi, A. Sami, J. Akhondali, F. Khomh, G. Uddin, and A. K. Motlagh, “An empirical study of c++ vulnerabilities in crowd-sourced code examples,” IEEE Trans. Softw. Eng. , vol. 48, no. 5, p. 1497–1514, may 2022. [Online]. Available: https://doi.org/10.1109/TSE.2020.3023664
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba, “Evaluating large language models trained on code,” 2021
2021
Earlier work this paper cites.
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton, “Program synthesis with large language models,” 2021
2021
Earlier work this paper cites.
A. Nikanjam, M. M. Morovati, F. Khomh, and H. Ben Braiek, “Faults in deep reinforcement learning programs: a taxonomy and a detection approach,” Automated Software Engineering , vol. 29, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. Imai, “Is github copilot a substitute for human pair-programming? an empirical study,” in Proceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings , 2022, pp. 319–321
2022
Earlier work this paper cites.
C. Bird, D. Ford, T. Zimmermann, N. Forsgren, E. Kalliamvakou, T. Lowdermilk, and I. Gazit, “Taking flight with copilot: Early insights and opportunities of ai-powered pair-programming tools,” Queue , vol. 20, no. 6, pp. 35–57, 2022
2022
Earlier work this paper cites.
N. Nguyen and S. Nadi, “An empirical evaluation of github copilot’s code suggestions,” in Proceedings of the 19th International Conference on Mining Software Repositories , 2022, pp. 1–5
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
“gitindex,” 2023, https://githut.info/
2023
Later among the works it cites.
“tiobeindex,” 2023, https://www.tiobe.com/tiobe-index/
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
V. Guilherme and A. Vincenzi, “An initial investigation of chatgpt unit test generation capability,” in Proceedings of the 8th Brazilian Symposium on Systematic and Automated Software Testing , ser. SAST ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 15–24. [Online]. Available: https://doi.org/10.1145/3624032.3624035
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Wei, X. Wang, D. Schuurmans, M. Bosma, b. ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35. Curran Associates, Inc., 2022, pp. 24 824–24 837. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf
2022
Cited alongside, same era.
P. Vaithilingam, T. Zhang, and E. L. Glassman, “Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,” in CHI Conference on Human Factors in Computing Systems Extended Abstracts , 2022, pp. 1–7
2022
Cited alongside, same era.
“Google forms,” https://www.google.ca/forms/about/ , 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
M. Jin, S. Shahriar, M. Tufano, X. Shi, S. Lu, N. Sundaresan, and A. Svyatkovskiy, “Inferfix: End-to-end program repair with llms,” in Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering , ser. ESEC/FSE 2023. New York, NY, USA: Association for Computing Machinery, 2023, p. 1646–1656. [Online]. Available: https://doi.org/10.1145/3611643.3613892
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
Y. Tang, Z. Liu, Z. Zhou, and X. Luo, “Chatgpt vs sbst: A comparative assessment of unit test suite generation,” 2023
2023
Later among the works it cites.
G. L. Scoccia, “Exploring early adopters’ perceptions of chatgpt as a code generation tool,” in 2023 38th IEEE/ACM International Conference on Automated Software Engineering Workshops (ASEW) , 2023, pp. 88–93
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Li, C. Wang, Z. Liu, H. Wang, D. Chen, S. Wang, and C. Gao, “Cctest: Testing and repairing code completion systems,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) , 2023, pp. 1238–1250
2023
Later among the works it cites.
S. Moon, Y. Song, H. Chae, D. Kang, T. Kwon, K. T. iunn Ong, S. won Hwang, and J. Yeo, “Coffee: Boost your code llms by fixing bugs with feedback,” 2023
2023
Later among the works it cites.
“Leetcode contest,” https://leetcode.com/contest/ , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Du, M. Liu, K. Wang, H. Wang, J. Liu, Y. Chen, J. Feng, C. Sha, X. Peng, and Y. Lou, “Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Ji, P. Ma, Z. Li, and S. Wang, “Benchmarking and explaining large language model-based code generation: A causality-centric approach,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Jesse, T. Ahmed, P. T. Devanbu, and E. Morgan, “Large language models and simple, stupid bugs,” in 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR) . Los Alamitos, CA, USA: IEEE Computer Society, may 2023, pp. 563–575. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/MSR59073.2023.00082
2023
Later among the works it cites.
J. Bogatinovski and O. Kao, “Auto-logging: AI-centred logging instrumentation,” in 2023 IEEE/ACM 45th International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER) . IEEE, 2023, pp. 95–100
2023
Later among the works it cites.
2024
Closest in time.
F. Tambon, A. Nikanjam, L. An, F. Khomh, and G. Antoniol, “Silent bugs in deep learning frameworks: an empirical study of keras and tensorflow,” Empirical Software Engineering , vol. 29, no. 1, p. 10, 2024
2024
Closest in time.
C. Liu, S. D. Zhang, and R. Jabbarvand, “Codemind: A framework to challenge large language models for code reasoning,” 2024
2024
Closest in time.
D. Huang, J. M. Zhang, Y. Qing, and H. Cui, “Effibench: Benchmarking the efficiency of automatically generated code,” 2024
2024
Closest in time.
M. M. Morovati, A. Nikanjam, F. Tambon, F. Khomh, and Z. M. Jiang, “Bug characterization in machine learning-based systems,” Empirical Software Engineering , vol. 29, no. 1, p. 14, 2024
2024
Closest in time.