Fetching the paper…
Reading the bibliography…
Existing methods for code generation use code snippets as seed data, restricting the complexity and diversity of the synthesized data.
A complexity measure
McCabe, T. J · 1976
Earlier work this paper cites.
Elements of Software Science (Operating and programming systems series)
Halstead, M. H · 1977
Earlier work this paper cites.
On the "naturalness" of buggy code
Ray, B., Hellendoorn, V., Godhane, S., Tu, Z., Bacchelli, A., and Devanbu, P · 2016
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Sener, O. and Savarese, S · 2018
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Earlier work this paper cites.
OPT: Open Pre-trained Transformer Language Models
Zhang, S., Roller, S., Goyal, N., and Artetxe · 2022
Earlier work this paper cites.
Code alpaca: An instruction-following llama model for code generation
Chaudhary, S · 2023
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Liu, J., Xia, C. S., Wang, Y., and Zhang, L · 2023
Cited alongside, same era.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D · 2023
Cited alongside, same era.
Octopack: Instruction tuning code large language models
Muennighoff, N., Liu, Q., Zebaze, A., Zheng, Q., Hui, B., Zhuo, T. Y., Singh, S., Tang, X., Von Werra, L., and Longpre, S · 2023
Cited alongside, same era.
Rewriting the code: A simple method for large language model augmented code search
Li, H., Zhou, X., and Shen, Z · 2024
Later among the works it cites.
Fullstack bench: Evaluating llms as full stack coder
Liu, S., Zhu, H., Liu, J., Xin, S., Li, A., Long, R., Chen, L., Yang, J., Xia, J., Peng, Z., et al · 2024
Later among the works it cites.
Starcoder 2 and the stack v2: The next generation
Lozhkov, A., Li, R., Allal, L. B., Cassano, F., Lamy-Poirier, J., Tazi, N., Tang, A., Pykhtar, D., Liu, J., Wei, Y., et al · 2024
Later among the works it cites.
Genetic instruct: Scaling up synthetic generation of coding instructions for large language models
Majumdar, S., Noroozi, V., Narenthiran, S., Ficek, A., Balam, J., and Ginsburg, B · 2024
Later among the works it cites.
A survey of neural code intelligence: Paradigms, advances and beyond
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GPT-4 Technical Report
OpenAI · 2023
Cited alongside, same era.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al · 2023
Cited alongside, same era.
Data augmentation using LLMs: Data perspectives, learning paradigms and challenges
Ding, B., Qin, C., Zhao, R., Luo, T., Li, X., Chen, G., Xia, W., Hu, J., Luu, A. T., and Joty, S · 2024
Cited alongside, same era.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
Guo, D., Zhu, Q., Yang, D., Xie, Z., Dong, K., Zhang, W., Chen, G., Bi, X., Wu, Y., Li, Y., et al · 2024
Cited alongside, same era.
Qwen2. 5-coder technical report
Hui, B., Yang, J., Cui, Z., Yang, J., Liu, D., Zhang, L., Liu, T., Zhang, J., Yu, B., Lu, K., et al · 2024
Cited alongside, same era.
Llm-based and retrieval-augmented control code generation
Koziolek, H., Grüner, S., Hark, R., Ashiwal, V., Linsbauer, S., and Eskandani, N · 2024
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H
Cited in the paper.
Sun, Q., Chen, Z., Xu, F., Cheng, K., Ma, C., Yin, Z., Wang, J., Han, C., Zhu, R., Yuan, S., et al · 2024
Later among the works it cites.
Top leaderboard ranking = top coding proficiency, always? evoeval: Evolving coding benchmarks via LLM
Xia, C. S., Deng, Y., and ZHANG, L · 2024
Later among the works it cites.
Opencodeinterpreter: Integrating code generation with execution and refinement
Zheng, T., Zhang, G., Shen, T., Liu, X., Lin, B. Y., Fu, J., Chen, W., and Yue, X · 2024
Later among the works it cites.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence
Zhu, Q., Guo, D., Shao, Z., Yang, D., Wang, P., Xu, R., Wu, Y., Li, Y., Gao, H., Ma, S., et al · 2024
Later among the works it cites.
Unigencoder: Merging seq2seq and seq2tree paradigms for unified code generation
Shao, L., Yan, Y., Poshyvanyk, D., and Su, J · 2025
Closest in time.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Zhuo, T. Y., Chien, V. M., Chim, J., Hu, H., Yu, W., Widyasari, R., Yusuf, I. N. B., Zhan, H., He, J., Paul, I., Brunner, S., GONG, C., Hoang, J., Zebaze, A. R., Hong, X., Li, W.-D., Kaddour, J., Xu, M., Zhang, Z., Yadav, P., Jain, N., Gu, A., Cheng, Z., Liu, J., Liu, Q., Wang, Z., Lo, D., Hui, B., Muennighoff, N., Fried, D., Du, X., de Vries, H., and Werra, L. V · 2025
Closest in time.