Fetching the paper…
Reading the bibliography…
Code quality evaluation involves scoring generated code quality based on a reference code for a specific problem statement.
M. G. Kendall, “A new measure of rank correlation,” Biometrika , vol. 30, no. 1/2, pp. 81–93, 1938
1938
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
L. C. ROUGE, “A package for automatic evaluation of summaries,” in Proceedings of Workshop on Text Summarization of ACL, Spain , vol. 5, 2004
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
I. Cohen, Y. Huang, J. Chen, J. Benesty, J. Benesty, J. Chen, Y. Huang, and I. Cohen, “Pearson correlation coefficient,” Noise reduction in speech processing , pp. 1–4, 2009
2009
Earlier work this paper cites.
M. Popović, “chrf: character n-gram f-score for automatic mt evaluation,” in Proceedings of the tenth workshop on statistical machine translation , 2015, pp. 392–395
2015
Earlier work this paper cites.
A. Joshi, S. Kale, S. Chandel, and D. K. Pal, “Likert scale: Explored and explained,” British journal of applied science & technology , vol. 7, no. 4, pp. 396–403, 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Yin, B. Deng, E. Chen, B. Vasilescu, and G. Neubig, “Learning to mine aligned code and natural language pairs from stack overflow,” in Proceedings of the 15th international conference on mining software repositories , 2018, pp. 476–486
2018
Earlier work this paper cites.
S. Kulal, P. Pasupat, K. Chandra, M. Lee, O. Padon, A. Aiken, and P. S. Liang, “Spoc: Search-based pseudocode to code,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
N. Tran, H. Tran, S. Nguyen, H. Nguyen, and T. Nguyen, “Does bleu score work for code migration?” in 2019 IEEE/ACM 27th International Conference on Program Comprehension (ICPC) . IEEE, 2019, pp. 165–176
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
P. R. Nair, “Increasing employability of indian engineering graduates through experiential learning programs and competitive programming: Case study,” Procedia Computer Science , vol. 172, pp. 831–837, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
B. Roziere, M.-A. Lachaux, L. Chanussot, and G. Lample, “Unsupervised translation of programming languages,” Advances in Neural Information Processing Systems , vol. 33, pp. 20 601–20 611, 2020
2020
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2023
Later among the works it cites.
M. Evtikhiev, E. Bogomolov, Y. Sokolov, and T. Bryksin, “Out of the bleu: how should we assess quality of the code generation models?” Journal of Systems and Software , vol. 203, p. 111741, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
H. Le, Y. Wang, A. D. Gotmare, S. Savarese, and S. C. H. Hoi, “Coderl: Mastering code generation through pretrained models and deep reinforcement learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 21 314–21 328, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 24 824–24 837, 2022
2022
Cited alongside, same era.
OpenAI., “Openai gpt-3.5 turbo,” https://platform.openai.com/docs/guides/text-generation/chat-completions-api , 2022
2022
Cited alongside, same era.
D. Das, N. S. Mathews, and S. Chimalakonda, “Exploring security vulnerabilities in competitive programming: An empirical study,” in Proceedings of the 26th International Conference on Evaluation and Assessment in Software Engineering , 2022, pp. 110–119
2022
Cited alongside, same era.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” Advances in neural information processing systems , vol. 35, pp. 22 199–22 213, 2022
2022
Cited alongside, same era.
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys , vol. 55, no. 12, pp. 1–38, 2023
2023
Later among the works it cites.
T. Hu, Z. Xu, Y. Fang, Y. Wu, B. Yuan, D. Zou, and H. Jin, “Fine-grained code clone detection with block-based splitting of abstract syntax tree,” in Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis , 2023, pp. 89–100
2023
Later among the works it cites.
E. Shi, Y. Wang, L. Du, H. Zhang, S. Han, D. Zhang, and H. Sun, “Cocoast: Representing source code via hierarchical splitting and reconstruction of abstract syntax trees,” Empirical Software Engineering , vol. 28, no. 6, pp. 1–41, 2023
2023
Later among the works it cites.
Y. Choi, H. Kim, and J.-H. Lee, “Blocsum: Block scope-based source code summarization via shared block representation,” in Findings of the Association for Computational Linguistics: ACL 2023 , 2023, pp. 11 427–11 441
2023
Later among the works it cites.
F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, A. Guha, M. Greenberg, and A. Jangda, “Multipl-e: A scalable and polyglot approach to benchmarking neural code generation,” IEEE Transactions on Software Engineering , vol. 49, no. 7, pp. 3675–3691, 2023
2023
Later among the works it cites.
OpenAI, “Gpt-4: Language models at scale,” https://openai.com/gpt-4/ , 2023
2023
Later among the works it cites.
OpenAI., “Openai gpt-4 turbo,” https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turbo , 2023
2023
Later among the works it cites.
S. Lukasczyk, F. Kroiß, and G. Fraser, “An empirical study of automated unit test generation for python,” Empirical Software Engineering , vol. 28, no. 2, p. 36, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.