Fetching the paper…
Reading the bibliography…
Retrieval-augmented generation (RAG) has increasingly shown its power in extending large language models' (LLMs') capability beyond their pre-trained knowledge.
Building More Usable APIs
Samuel G. McLellan, Alvin W. Roesler, Joseph T. Tempest, and Clay Spinuzzi. 1998 · 1998
Earlier work this paper cites.
Automated identification of parameter mismatches in web applications. In Proceedings of the 16th ACM SIGSOFT International Symposium on Foundations of Software Engineering, 2008, Atlanta, Georgia, USA, November 9-14, 2008 . ACM, 181–191
William G. J. Halfond and Alessandro Orso. 2008 · 2008
Earlier work this paper cites.
The Probabilistic Relevance Framework: BM25 and Beyond
Stephen E. Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
What Makes APIs Hard to Learn? Answers from Developers
Martin P. Robillard. 2009 · 2009
Earlier work this paper cites.
A field study of API learning obstacles
Martin P. Robillard and Robert DeLine. 2011 · 2011
Earlier work this paper cites.
Detecting API documentation errors. In Proceedings of the 2013 ACM SIGPLAN International Conference on Object Oriented Programming Systems Languages & Applications, OOPSLA 2013, part of SPLASH 2013, Indianapolis, IN, USA, October 26-31, 2013 . ACM, 803–816
Hao Zhong and Zhendong Su. 2013 · 2013
Earlier work this paper cites.
Automatic Detection and Update Suggestion for Outdated API Names in Documentation
Seonah Lee, Rongxin Wu, Shing-Chi Cheung, and Sungwon Kang. 2021 · 2019
Earlier work this paper cites.
Software documentation: the practitioners’ perspective. In ICSE ’20: 42nd International Conference on Software Engineering, Seoul, South Korea, 27 June - 19 July, 2020 . ACM, 590–601
Emad Aghajani, Csaba Nagy, Mario Linares-Vásquez, Laura Moreno, Gabriele Bavota, Michele Lanza, and David C. Shepherd. 2020 · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
Evaluating Large Language Models Trained on Code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Earlier work this paper cites.
Ivy: Templated Deep Learning for Inter-Framework Portability
Daniel Lenton, Fabio Pardo, Fabian Falck, Stephen James, and Ronald Clark. 2021 · 2021
Earlier work this paper cites.
Cocomic: Code completion by jointly modeling in-file and cross-file context
Yangruibo Ding, Zijian Wang, Wasi Uddin Ahmad, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, and Bing Xiang. 2022 · 2022
Earlier work this paper cites.
Text and Code Embeddings by Contrastive Pre-Training
Arvind Neelakantan and Tao Xu et al. 2022 · 2022
Earlier work this paper cites.
Enhancing Chat Language Models by Scaling High-quality Instructional Conversations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023 . Association for Computational Linguistics, 3029–3051
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023 · 2023
Earlier work this paper cites.
Efficient Text-to-Code Retrieval with Cascaded Fast and Slow Transformer Models. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 388–400
Akhilesh Deepak Gotmare, Junnan Li, Shafiq Joty, and Steven CH Hoi. 2023 · 2023
Earlier work this paper cites.
DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol. 202) . PMLR, 18319–18345
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-Tau Yih, Daniel Fried, Sida I. Wang, and Tao Yu. 2023 · 2023
Earlier work this paper cites.
Large language model-aware in-context learning for code generation
Jia Li, Chongyang Tao, Jia Li, Ge Li, Zhi Jin, Huangzhao Zhang, Zheng Fang, and Fang Liu. 2023a · 2023
Earlier work this paper cites.
Towards General Text Embeddings with Multi-stage Contrastive Learning
Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023b · 2023
Earlier work this paper cites.
MTEB: Massive Text Embedding Benchmark. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May 2-6, 2023 . Association for Computational Linguistics, 2006–2029
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2023 · 2023
Cited alongside, same era.
Fine-tuning or retrieval? comparing knowledge injection in llms
Oded Ovadia, Menachem Brief, Moshik Mishaeli, and Oren Elisha. 2023 · 2023
Cited alongside, same era.
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023a · 2023
Cited alongside, same era.
Retrieval-Augmented Test Generation: How Far Are We?
Jiho Shin, Reem Aleithan, Hadi Hemmati, and Song Wang. 2024 · 2024
Later among the works it cites.
Arks: Active retrieval in knowledge soup for code generation
Hongjin Su, Shuyang Jiang, Yuhang Lai, Haoyuan Wu, Boao Shi, Che Liu, Qian Liu, and Tao Yu. 2024a · 2024
Later among the works it cites.
EvoR: Evolving Retrieval for Code Generation. In Findings of the Association for Computational Linguistics: EMNLP 2024 . 2538–2554
Hongjin Su, Shuyang Jiang, Yuhang Lai, Haoyuan Wu, Boao Shi, Che Liu, Qian Liu, and Tao Yu. 2024b · 2024
Later among the works it cites.
Prompt-based code completion via multi-retrieval augmented generation
Hanzhuo Tan, Qi Luo, Ling Jiang, Zizheng Zhan, Jing Li, Haotian Zhang, and Yuqun Zhang. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-Instruct: Aligning Language Models with Self-Generated Instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 . Association for Computational Linguistics, 13484–13508
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023b · 2023
Cited alongside, same era.
Retrieval meets long context large language models. In The Twelfth International Conference on Learning Representations
Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, and Bryan Catanzaro. 2023 · 2023
Cited alongside, same era.
Repocoder: Repository-level code completion through iterative retrieval and generation
Fengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen. 2023a · 2023
Cited alongside, same era.
Syntax-aware retrieval augmented code generation. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 1291–1302
Xiangyu Zhang, Yu Zhou, Guang Yang, and Taolue Chen. 2023b · 2023
Cited alongside, same era.
Generating a Low-code Complete Workflow via Task Decomposition and RAG
Orlando Marquez Ayala and Patrice Béchard. 2024 · 2024
Cited alongside, same era.
Benchmarking Large Language Models in Retrieval-Augmented Generation. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada . AAAI Press, 17754–17762
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024a · 2024
Cited alongside, same era.
Dense x retrieval: What retrieval granularity should we use?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 15159–15177
Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Hongming Zhang, and Dong Yu. 2024b · 2024
Cited alongside, same era.
RAR: Retrieval-augmented retrieval for code generation in low resource languages. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024 , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, 21506–21515
Avik Dutta, Mukul Singh, Gust Verbruggen, Sumit Gulwani, and Vu Le. 2024 · 2024
Cited alongside, same era.
Exploring Demonstration Retrievers in RAG for Coding Tasks: Yeas and Nays!
Pengfei He, Shaowei Wang, Shaiful Chowdhury, and Tse-Hsun Chen. 2024 · 2024
Cited alongside, same era.
Chong Wang, Kaifeng Huang, Jian Zhang, Yebo Feng, Lyuye Zhang, Yang Liu, and Xin Peng. 2024b · 2024
Later among the works it cites.
Do Advanced Language Models Eliminate the Need for Prompt Engineering in Software Engineering?
Guoqing Wang, Zeyu Sun, Zhihao Gong, Sixiang Ye, Yizhou Chen, Yifan Zhao, Qingyuan Liang, and Dan Hao. 2024c · 2024
Later among the works it cites.
Searching for best practices in retrieval-augmented generation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 17716–17736
Xiaohua Wang, Zhenghua Wang, Xuan Gao, Feiran Zhang, Yixin Wu, Zhibo Xu, Tianyuan Shi, Zhengyuan Wang, Shizheng Li, Qi Qian, et al · 2024
Later among the works it cites.
CodeRAG-Bench: Can Retrieval Augment Code Generation?
Zora Zhiruo Wang, Akari Asai, Xinyan Velocity Yu, Frank F. Xu, Yiqing Xie, Graham Neubig, and Daniel Fried. 2024a · 2024
Later among the works it cites.
Programming by Example Made Easy
Jiarong Wu, Lili Wei, Yanyan Jiang, Shing-Chi Cheung, Luyao Ren, and Chang Xu. 2024 · 2024
Later among the works it cites.
MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic Testing. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, ASE 2024, Sacramento, CA, USA, October 27 - November 1, 2024 , Vladimir Filkov, Baishakhi Ray, and Minghui Zhou (Eds.). ACM, 557–569
Congying Xu, Songqiang Chen, Jiarong Wu, Shing-Chi Cheung, Valerio Terragni, Hengcheng Zhu, and Jialun Cao. 2024 · 2024
Later among the works it cites.
DroidCoder: Enhanced Android Code Completion with Context-Enriched Retrieval-Augmented Generation. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, ASE 2024, Sacramento, CA, USA, October 27 - November 1, 2024 . ACM, 681–693
Xinran Yu, Chun Li, Minxue Pan, and Xuandong Li. 2024 · 2024
Later among the works it cites.
Towards Understanding Retrieval Accuracy and Prompt Quality in RAG Systems
Shengming Zhao, Yuheng Huang, Jiayang Song, Zhijie Wang, Chengcheng Wan, and Lei Ma. 2024 · 2024
Later among the works it cites.
DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code Generation
Qiming Zhu, Jialun Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun, and Shing-Chi Cheung. 2024 · 2024
Later among the works it cites.
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Terry Yue Zhuo and Minh Chien Vu et al. 2024 · 2024
Later among the works it cites.
Jialun Cao, Yaojie Lu, Meiziniu Li, Haoyang Ma, Haokun Li, Mengda He, Cheng Wen, Le Sun, Hongyu Zhang, Shengchao Qin, Shing-Chi Cheung, and Cong Tian. 2025 · 2025
Closest in time.
An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities
Zezhou Yang, Sirong Chen, Cuiyun Gao, Zhenhao Li, Xing Hu, Kui Liu, and Xin Xia. 2025 · 2025
Closest in time.
Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024 . Association for Computational Linguistics, 2053–2065
Rui Li, Qi Liu, Liyang He, Zheng Zhang, Hao Zhang, Shengyu Ye, Junyu Lu, and Zhenya Huang. 2024 · 2065
Closest in time.