Fetching the paper…
Reading the bibliography…
To evaluate code large language models (LLMs), research has relied on a few small manually curated benchmarks, such as HumanEval and MBPP, which represent a narrow part of the real-world software domains.
Property-based testing: a new approach to testing for assurance
Fink, G. and Bishop, M · 1997
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Automatic test generation: A use case driven approach
Nebut, C., Fleurey, F., Le Traon, Y., and Jezequel, J.-M · 2006
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
A syntactic neural model for general-purpose code generation
Yin, P. and Neubig, G · 2017
Earlier work this paper cites.
Understanding back-translation at scale
Edunov, S., Ott, M., Auli, M., and Grangier, D · 2018
Earlier work this paper cites.
JuICe: A large scale distantly supervised dataset for open domain context-based code generation
Agashe, R., Iyer, S., and Zettlemoyer, L · 2019
Earlier work this paper cites.
Data augmentation using back-translation for context-aware neural machine translation
Sugiyama, A. and Yoshinaga, N · 2019
Earlier work this paper cites.
Code generation as a dual task of code summarization
Wei, B., Li, G., Xia, X., Fu, Z., and Jin, Z · 2019
Earlier work this paper cites.
BERTScore: Evaluating text generation with BERT
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2019
Earlier work this paper cites.
CodeBLEU: a method for automatic evaluation of code synthesis
Ren, S., Guo, D., Lu, S., Zhou, L., Liu, S., Tang, D., Sundaresan, N., Zhou, M., Blanco, A., and Ma, S · 2020
Cited alongside, same era.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Cited alongside, same era.
On multi-modal learning of editing source code
Chakraborty, S. and Ray, B · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Cited alongside, same era.
DS-1000: A natural and reliable benchmark for data science code generation
Lai, Y., Li, C., Wang, Y., Zhang, T., Zhong, R., Zettlemoyer, L., Yih, W.-t., Fried, D., Wang, S., and Yu, T · 2023
Later among the works it cites.
Beyond accuracy: Evaluating self-consistency of code large language models with IdentityChain
Min, M. J., Ding, Y., Buratti, L., Pujar, S., Kaiser, G., Jana, S., and Ray, B · 2023
Later among the works it cites.
OctoPack: Instruction tuning code large language models
Muennighoff, N., Liu, Q., Zebaze, A., Zheng, Q., Hui, B., Zhuo, T. Y., Singh, S., Tang, X., Von Werra, L., and Longpre, S · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team Gemini, Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., et al · 2021
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D · 2022
Cited alongside, same era.
Natural language to code generation in interactive data science notebooks
Yin, P., Li, W.-D., Xiao, K., Rao, A., Wen, Y., Shi, K., Howland, J., Bailey, P., Catasta, M., Michalewski, H., et al · 2022
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Cited alongside, same era.
Faithfulness tests for natural language explanations
Atanasova, P., Camburu, O.-M., Lioma, C., Lukasiewicz, T., Simonsen, J. G., and Augenstein, I · 2023
Cited alongside, same era.
SWE-bench: Can language models resolve real-world GitHub issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2023
Cited alongside, same era.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al
Cited in the paper.
Automating code review activities by large-scale pre-training
Li, Z., Lu, S., Guo, D., Duan, N., Jannu, S., Jenks, G., Majumder, D., Green, J., Svyatkovskiy, A., Fu, S., et al
Cited in the paper.
Zhou, S., Alon, U., Agarwal, S., and Neubig, G · 2023
Later among the works it cites.
CrossCodeEval: A diverse and multilingual benchmark for cross-file code completion
Ding, Y., Wang, Z., Ahmad, W., Ding, H., Tan, M., Jain, N., Ramanathan, M. K., Nallapati, R., Bhatia, P., Roth, D., et al · 2024
Closest in time.
DeepSeek-Coder: When the large language model meets programming – the rise of code intelligence
Guo, D., Zhu, Q., Yang, D., Xie, Z., Dong, K., Zhang, W., Chen, G., Bi, X., Wu, Y., Li, Y., et al · 2024
Closest in time.
CodeMind: A framework to challenge large language models for code reasoning
Liu, C., Zhang, S. D., and Jabbarvand, R · 2024
Closest in time.
StarCoder 2 and The Stack v2: The next generation, 2024
Lozhkov, A., Li, R., Allal, L. B., Cassano, F., Lamy-Poirier, J., Tazi, N., Tang, A., Pykhtar, D., Liu, J., Wei, Y., Liu, T., Tian, M., Kocetkov, D., Zucker, A., Belkada, Y., Wang, Z., Liu, Q., Abulkhanov, D., Paul, I., Li, Z., Li, W.-D., Risdal, M., Li, J., Zhu, J., Zhuo, T. Y., Zheltonozhskii, E., Dade, N. O. O., Yu, W., Krauß, L., Jain, N., Su, Y., He, X., Dey, M., Abati, E., Chai, Y., Muennighoff, N., Tang, X., Oblokulov, M., Akiki, C., Marone, M., Mou, C., Mishra, M., Gu, A., Hui, B., Dao, T., Zebaze, A., Dehaene, O., Patry, N., Xu, C., McAuley, J., Hu, H., Scholak, T., Paquet, S., Robinson, J., Anderson, C. J., Chapados, N., Patwary, M., Tajbakhsh, N., Jernite, Y., Ferrandis, C. M., Zhang, L., Hughes, S., Wolf, T., Guha, A., von Werra, L., and de Vries, H · 2024
Closest in time.