Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have exhibited exceptional performance in software engineering yet face challenges in adapting to continually evolving code knowledge, particularly regarding the frequent updates of third-party library APIs.
Learning string-edit distance
Ristad, E. and Yianilos, P · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Codebleu: a method for automatic evaluation of code synthesis
Ren, S., Guo, D., Lu, S., Zhou, L., Liu, S., Tang, D., Sundaresan, N., Zhou, M., Blanco, A., and Ma, S · 2009
Earlier work this paper cites.
Long short-term memory
Graves, A. and Graves, A · 2012
Earlier work this paper cites.
How do software engineers understand code changes? an exploratory study in industry
Tao, Y., Dang, Y., Xie, T., Zhang, D., and Kim, S · 2012
Earlier work this paper cites.
Gated feedback recurrent neural networks
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
A convolutional attention network for extreme summarization of source code
Allamanis, M., Peng, H., and Sutton, C · 2016
Earlier work this paper cites.
Deep api learning
Gu, X., Zhang, H., Zhang, D., and Kim, S · 2016
Earlier work this paper cites.
Summarizing source code using a neural attention model
Iyer, S., Konstas, I., Cheung, A., and Zettlemoyer, L · 2016
Earlier work this paper cites.
Convolutional neural networks over tree structures for programming language processing
Mou, L., Li, G., Zhang, L., Wang, T., and Jin, Z · 2016
Earlier work this paper cites.
Exploring api embedding for api usages and applications
Nguyen, T. D., Nguyen, A. T., Phan, H. D., and Nguyen, T. N · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Learning to represent programs with graphs
Allamanis, M., Brockschmidt, M., and Khademi, M · 2018
Earlier work this paper cites.
code2seq: Generating sequences from structured representations of code
Alon, U., Brody, S., Levy, O., and Yahav, E · 2018
Earlier work this paper cites.
Deep code search
Gu, X., Zhang, H., and Kim, S · 2018
Earlier work this paper cites.
Improving automatic source code summarization via deep reinforcement learning
Wan, Y., Zhao, Z., Yang, M., Xu, G., Ying, H., Wu, J., and Yu, P. S · 2018
Earlier work this paper cites.
Convolutional neural networks: an overview and application in radiology
Yamashita, R., Nishio, M., Do, R. K. G., and Togashi, K · 2018
Earlier work this paper cites.
code2vec: Learning distributed representations of code
Alon, U., Zilberstein, M., Levy, O., and Yahav, E · 2019
Earlier work this paper cites.
Multi-modal attention network learning for semantic source code retrieval
Wan, Y., Shu, J., Sui, Y., Xu, G., Zhao, Z., Wu, J., and Yu, P · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Ir2vec: Llvm ir based scalable program embeddings
VenkataKeerthy, S., Aggarwal, R., Jain, S., Desarkar, M. S., Upadrasta, R., and Srikant, Y · 2020
Earlier work this paper cites.
Reinforcement-learning-guided source code summarization using hierarchical attention
Wang, W., Zhang, Y., Sui, Y., Wan, Y., Zhao, Z., Wu, J., Philip, S. Y., and Xu, G · 2020
Earlier work this paper cites.
How do python framework apis evolve? an exploratory study
Zhang, Z., Zhu, H., Wen, M., Tao, Y., Liu, Y., and Xiong, Y · 2020
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Lu, S., Guo, D., Ren, S., Huang, J., Svyatkovskiy, A., Blanco, A., Clement, C. B., Drain, D., Jiang, D., Tang, D., Li, G., Zhou, L., Shou, L., Zhou, L., Tufano, M., Gong, M., Zhou, M., Duan, N., Sundaresan, N., Deng, S. K., Fu, S., and Liu, S · 2021
Earlier work this paper cites.
How could neural networks understand programs?
Peng, D., Zheng, S., Li, Y., Ke, G., He, D., and Liu, T.-Y · 2021
Cited alongside, same era.
Synthetic data augmentation for zero-shot cross-lingual question answering
Riabi, A., Scialom, T., Keraron, R., Sagot, B., Seddah, D., and Staiano, J · 2021
Cited alongside, same era.
Generating datasets with pretrained language models
Schick, T. and Schütze, H · 2021
Cited alongside, same era.
A framework for the evaluation of code generation models
Ben Allal, L., Muennighoff, N., Kumar Umapathi, L., Lipkin, B., and von Werra, L · 2022
Cited alongside, same era.
Cross-language binary-source code matching with intermediate representations
Gui, Y., Wan, Y., Zhang, H., Huang, H., Sui, Y., Xu, G., Shao, Z., and Jin, H · 2022
Cited alongside, same era.
ORPO: Monolithic preference optimization without reference model
Hong, J., Lee, N., and Thorne, J · 2024
Later among the works it cites.
Trustllm: Trustworthiness in large language models
Huang, Y., Sun, L., Wang, H., Wu, S., Zhang, Q., Li, Y., Gao, C., Huang, Y., Lyu, W., Zhang, Y., et al · 2024
Later among the works it cites.
Qwen2.5-coder technical report
Hui, B., Yang, J., Cui, Z., Yang, J., Liu, D., Zhang, L., Liu, T., Zhang, J., Yu, B., Dang, K., et al · 2024
Later among the works it cites.
A survey on large language models for code generation
Jiang, J., Wang, F., Shen, J., Kim, S., and Kim, S · 2024
Later among the works it cites.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chung, J. J. Y., Kamar, E., and Amershi, S · 2023
Cited alongside, same era.
Incoder: A generative model for code infilling and synthesis
Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., Yih, S., Zettlemoyer, L., and Lewis, M · 2023
Cited alongside, same era.
Aging with grace: Lifelong model editing with discrete key-value adaptors
Hartvigsen, T., Sankaranarayanan, S., Palangi, H., Kim, Y., and Ghassemi, M · 2023
Cited alongside, same era.
Faithful persona-based conversational dataset generation with large language models
Jandaghi, P., Sheng, X., Bai, X., Pujara, J., and Sidahmed, H · 2023
Cited alongside, same era.
Starcoder: may the source be with you!
Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., Akiki, C., Li, J., Chim, J., Liu, Q., Zheltonozhskii, E., Zhuo, T. Y., Wang, T., Dehaene, O., Davaadorj, M., Lamy-Poirier, J., Monteiro, J., Shliazhko, O., Gontier, N., Meade, N., Zebaze, A., Yee, M., Umapathi, L. K., Zhu, J., Lipkin, B., Oblokulov, M., Wang, Z., V, R. M., Stillerman, J. T., Patel, S. S., Abulkhanov, D., Zocca, M., Dey, M., Zhang, Z., Fahmy, N., Bhattacharyya, U., Yu, W., Singh, S., Luccioni, S., Villegas, P., Kunakov, M., Zhdanov, F., Romero, M., Lee, T., Timor, N., Ding, J., Schlesinger, C., Schoelkopf, H., Ebert, J., Dao, T., Mishra, M., Gu, A., Robinson, J., Anderson, C. J., Dolan-Gavitt, B., Contractor, D., Reddy, S., Fried, D., Bahdanau, D., Jernite, Y., Ferrandis, C. M., Hughes, S., Wolf, T., Guha, A., von Werra, L., and de Vries, H · 2023
Cited alongside, same era.
Chatgpt: A conversational ai model, 2023
OpenAI · 2023
Cited alongside, same era.
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Cited alongside, same era.
Later among the works it cites.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D · 2024
Later among the works it cites.
Hello GPT-4o
OpenAI · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Qwen Team · 2024
Later among the works it cites.
Sifting through the chaff: On utilizing execution feedback for ranking the generated code candidates
Sun, Z., Wan, Y., Li, J., Zhang, H., Jin, Z., Li, G., and Lyu, C · 2024
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2024
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., Silver, D., Johnson, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., Lillicrap, T., and Angeliki Lazaridou, …, O. V · 2024
Later among the works it cites.
Deep learning for code intelligence: Survey, benchmark and toolkit
Wan, Y., Bi, Z., He, Y., Zhang, J., Zhang, H., Sui, Y., Xu, G., Jin, H., and Yu, P · 2024
Later among the works it cites.
Unigen: A unified framework for textual dataset generation using large language models
Wu, S., Huang, Y., Gao, C., Chen, D., Zhang, Q., Wan, Y., Zhou, T., Zhang, X., Gao, J., Xiao, C., et al · 2024
Later among the works it cites.
Justice or prejudice? quantifying biases in llm-as-a-judge
Ye, J., Wang, Y., Huang, Y., Chen, D., Zhang, Q., Moniz, N., Gao, T., Geyer, W., Huang, C., Chen, P.-Y., et al · 2024
Later among the works it cites.
LLM-as-a-coauthor: Can mixed human-written and machine-generated text be detected?
Zhang, Q., Gao, C., Chen, D., Huang, Y., Huang, Y., Sun, Z., Zhang, S., Li, W., Fu, Z., Wan, Y., and Sun, L · 2024
Later among the works it cites.
GUI-world: A GUI-oriented dataset for multimodal LLM-based agents
Chen, D., Huang, Y., Wu, S., Tang, J., Zhou, H., Zhang, Q., He, Z., Bai, Y., Gao, C., Chen, L., Li, Y., Wang, C., Yu, Y., Zhou, T., Li, Z., Gui, Y., Wan, Y., Zhou, P., Gao, J., and Sun, L · 2025
Closest in time.
Auggpt: Leveraging chatgpt for text data augmentation
Dai, H., Liu, Z., Liao, W., Huang, X., Cao, Y., Wu, Z., Zhao, L., Xu, S., Zeng, F., Liu, W., et al · 2025
Closest in time.
Livevqa: Live visual knowledge seeking
Fu, M., Peng, Y., Liu, B., Wan, Y., and Chen, D · 2025
Closest in time.
Github code search
GitHub · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Wikipedia in the era of llms: Evolution and risks
Huang, S., Xu, Y., Geng, M., Wan, Y., and Chen, D · 2025
Closest in time.
CodeMMLU: A multi-task benchmark for assessing code understanding capabilities of codeLLMs
Nguyen, D. M., Phan, T. C., Hai, N. L., Doan, T.-T., Nguyen, N. V., Pham, Q., and Bui, N. D. Q · 2025
Closest in time.
GPT-4 Turbo and GPT-4 documentation
OpenAI · 2025
Closest in time.
Judge anything: Mllm as a judge across any modality
Pu, S., Wang, Y., Chen, D., Chen, Y., Wang, G., Qin, Q., Zhang, Z., Zhang, Z., Zhou, Z., Gong, S., et al · 2025
Closest in time.
ast — abstract syntax trees
Python · 2025
Closest in time.
inspect — inspect live objects
Python · 2025
Closest in time.