Fetching the paper…
Reading the bibliography…
RAG (Retrieval Augmented Generation) allows LLMs (large language models) to generate better responses with external knowledge, but using more external knowledge often improves generation quality at the expense of response delay.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks, 2019
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Transformers: State-of-the-Art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Perric Cistac, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Earlier work this paper cites.
QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization
Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev · 2021
Earlier work this paper cites.
LangChain, October 2022
Harrison Chase · 2022
Earlier work this paper cites.
Text and code embeddings by contrastive pre-training, 2022
Arvind Neelakantan et al · 2022
Earlier work this paper cites.
Long document re-ranking with modular re-ranker
Luyu Gao and Jamie Callan · 2022
Earlier work this paper cites.
MuSiQue: Multihop questions via single-hop question composition
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal · 2022
Earlier work this paper cites.
Cohere: Cutting-edge gen ai, 2023
Cohere AI · 2023
Earlier work this paper cites.
Retrieval-based language models and applications
Akari Asai, Sewon Min, Zexuan Zhong, and Danqi Chen · 2023
Earlier work this paper cites.
Active retrieval augmented generation, 2023
Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Earlier work this paper cites.
Query rewriting for retrieval-augmented large language models, 2023
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan · 2023
Earlier work this paper cites.
Zero-shot listwise document reranking with a large language model, 2023
Xueguang Ma, Xinyu Zhang, Ronak Pradeep, and Jimmy Lin · 2023
Earlier work this paper cites.
Openai api, 2023
OpenAI · 2023
Earlier work this paper cites.
Stateful large language model serving with pensieve
Lingfan Yu and Jinyang Li · 2023
Earlier work this paper cites.
https://docs.llamaindex.ai/en/stable/examples/param_optimizer/param_optimizer/ , 2024
Hyperparameter Optimization for RAG · 2024
Earlier work this paper cites.
Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee · 2024
Earlier work this paper cites.
PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation
Jason et al. Ansel · 2024
Earlier work this paper cites.
Rag vs fine-tuning: Pipelines, tradeoffs, and a case study on agriculture, 2024
Angels Balaguer, Vinamra Benara, Renato Luiz de Freitas Cunha, Roberto de M. Estevão Filho, Todd Hendry, Daniel Holstein, Jennifer Marsman, Nick Mecklenburg, Sara Malvar, Leonardo O. Nunes, Rafael Padilha, Morris Sharp, Bruno Silva, Swati Sharma, Vijay Aski, and Ranveer Chandra · 2024
Earlier work this paper cites.
Do large language models need a content delivery network?, 2024
Yihua Cheng, Kuntai Du, Jiayi Yao, and Junchen Jiang · 2024
Earlier work this paper cites.
The power of noise: Redefining retrieval for rag systems
Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, and Fabrizio Silvestri · 2024
Earlier work this paper cites.
Apparate: Rethinking early exits to tame latency-throughput tensions in ml serving
Yinwei Dai, Rui Pan, Anand Iyer, Kai Li, and Ravi Netravali · 2024
Earlier work this paper cites.
Hybrid llm: Cost-efficient and quality-aware query routing, 2024
Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Ruhle, Laks V. S. Lakshmanan, and Ahmed Hassan Awadallah · 2024
Earlier work this paper cites.
Get more with less: Synthesizing recurrence with kv cache compression for efficient llm inference, 2024
Harry Dong, Xinyu Yang, Zhenyu Zhang, Zhangyang Wang, Yuejie Chi, and Beidi Chen · 2024
Earlier work this paper cites.
Domain-driven llm development: Insights into rag and fine-tuning practices
José Cassio dos Santos Junior, Rachel Hu, Richard Song, and Yunfei Bai · 2024
Earlier work this paper cites.
The faiss library
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou · 2024
Earlier work this paper cites.
From local to global: A graph rag approach to query-focused summarization, 2024
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, and Jonathan Larson · 2024
Earlier work this paper cites.
Mteb: Massive text embedding benchmark, 2024
Hugging Face · 2024
Earlier work this paper cites.
A survey on rag meeting llms: Towards retrieval-augmented large language models
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li · 2024
Earlier work this paper cites.
Autorag-hp: Automatic online hyper-parameter tuning for retrieval-augmented generation
Jia Fu, Xiaoting Qin, Fangkai Yang, Lu Wang, Jue Zhang, Qingwei Lin, Yubo Chen, Dongmei Zhang, Saravan Rajmohan, and Qi Zhang · 2024
Earlier work this paper cites.
Cost-efficient large language model serving for multi-turn conversations with cachedattention
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo · 2024
Earlier work this paper cites.
From rag to fabric: Lessons learned from building real-world rags at genaiic – part 1
Aude Genevay · 2024
Earlier work this paper cites.
A survey of confidence estimation and calibration in large language models
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych · 2024
Cited alongside, same era.
Prompt cache: Modular attention reuse for low-latency inference
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong · 2024
Cited alongside, same era.
Lightrag: Simple and fast retrieval-augmented generation, 2024
Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and Chao Huang · 2024
Cited alongside, same era.
Ruler: What’s the real context size of your long-context language models?, 2024
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg · 2024
Cited alongside, same era.
Memserve: Context caching for disaggregated llm serving with elastic memory pool, 2024
Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan · 2024
Cited alongside, same era.
Beyond text: Optimizing rag with multimodal inputs for industrial applications, 2024
Monica Riedler and Stefan Langer · 2024
Closest in time.
Ragchecker: A fine-grained framework for diagnosing retrieval-augmented generation, 2024
Dongyu Ru, Lin Qiu, Xiangkun Hu, Tianhang Zhang, Peng Shi, Shuaichen Chang, Cheng Jiayang, Cunxiang Wang, Shichao Sun, Huanyu Li, Zizhao Zhang, Binjie Wang, Jiarong Jiang, Tong He, Zhiguo Wang, Pengfei Liu, Yue Zhang, and Zheng Zhang · 2024
Closest in time.
Fast inference for augmented large language models, 2024
Rana Shahout, Cong Liang, Shiji Xin, Qianru Lao, Yong Cui, Minlan Yu, and Michael Mitzenmacher · 2024
Closest in time.
Scaling retrieval-based language models with a trillion-token datastore
Rulin Shao, Jacqueline He, Akari Asai, Weijia Shi, Tim Dettmers, Sewon Min, Luke Zettlemoyer, and Pang Wei Koh · 2024
Closest in time.
A methodology for evaluating rag systems: A case study on configuration dependency validation, 2024
Sebastian Simon, Alina Mailach, Johannes Dorn, and Norbert Siegmund · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Epic: Efficient position-independent context caching for serving large language models, 2024
Junhao Hu, Wenrui Huang, Haoyi Wang, Weidong Wang, Tiancheng Hu, Qin Zhang, Hao Feng, Xusheng Chen, Yizhou Shan, and Tao Xie · 2024
Cited alongside, same era.
Routerbench: A benchmark for multi-llm routing system, 2024
Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay · 2024
Cited alongside, same era.
Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity
Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong Park · 2024
Cited alongside, same era.
Piperag: Fast retrieval-augmented generation via algorithm-system co-design
Wenqi Jiang, Shuai Zhang, Boran Han, Jie Wang, Bernie Wang, and Tim Kraska · 2024
Cited alongside, same era.
Neo: Saving gpu memory crisis with cpu offloading for online llm inference, 2024
Xuanlin Jiang, Yang Zhou, Shiyi Cao, Ion Stoica, and Minlan Yu · 2024
Cited alongside, same era.
Ragcache: Efficient knowledge caching for retrieval-augmented generation, 2024
Chao Jin, Zili Zhang, Xuanlin Jiang, Fangyue Liu, Xin Liu, Xuanzhe Liu, and Xin Jin · 2024
Cited alongside, same era.
Compute or load kv cache? why not both?, 2024
Shuowei Jin, Xueshen Liu, Qingzhao Zhang, and Z. Morley Mao · 2024
Cited alongside, same era.
Powerinfer: Fast large language model serving with a consumer-grade gpu
Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen · 2024
Closest in time.
Preble: Efficient distributed prompt scheduling for llm serving
Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, and Yiying Zhang · 2024
Closest in time.
Teola: Towards end-to-end optimization of llm-based applications, 2024
Xin Tan, Yimin Jiang, Yitao Yang, and Hong Xu · 2024
Closest in time.
Mba-rag: a bandit approach for adaptive retrieval-augmented generation through question complexity, 2024
Xiaqiang Tang, Qiang Gao, Jian Li, Nan Du, Qi Li, and Sihong Xie · 2024
Closest in time.
Searching for best practices in retrieval-augmented generation
Xiaohua Wang, Zhenghua Wang, Xuan Gao, Feiran Zhang, Yixin Wu, Zhibo Xu, Tianyuan Shi, Zhengyuan Wang, Shizheng Li, Qi Qian, Ruicheng Yin, Changze Lv, Xiaoqing Zheng, and Xuanjing Huang · 2024
Closest in time.
Model tells you where to merge: Adaptive kv cache merging for llms on long-context tasks, 2024
Zheng Wang, Boxiao Jin, Zhongzhi Yu, and Minjia Zhang · 2024
Closest in time.
Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism
Bingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun, Xuanzhe Liu, and Xin Jin · 2024
Closest in time.
Weknow-rag: An adaptive approach for retrieval-augmented generation integrating web search and knowledge graphs, 2024
Weijian Xie, Xuefeng Liang, Yuhui Liu, Kaihua Ni, Hong Cheng, and Zetian Hu · 2024
Closest in time.
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi · 2024
Closest in time.
Cacheblend: Fast large language model serving for rag with cached knowledge fusion, 2024
Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang · 2024
Closest in time.
Pqcache: Product quantization-based kvcache for long context llm inference, 2024
Hailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu, Xupeng Miao, Xiaonan Nie, Weipeng Chen, and Bin Cui · 2024
Closest in time.
Caravan: Practical online learning of In-Network ML models with labeling agents
Qizheng Zhang, Ali Imran, Enkeleda Bardhi, Tushar Swamy, Nathan Zhang, Muhammad Shahbaz, and Kunle Olukotun · 2024
Closest in time.
Raft: Adapting language model to domain specific rag, 2024
Tianjun Zhang, Shishir G. Patil, Naman Jain, Sheng Shen, Matei Zaharia, Ion Stoica, and Joseph E. Gonzalez · 2024
Closest in time.
Optimizing LLM based retrieval augmented generation pipelines in the financial domain
Yiyun Zhao, Prateek Singh, Hanoz Bhathena, Bernardo Ramos, Aviral Joshi, Swaroop Gadiyaram, and Saket Sharma · 2024
Closest in time.
SGLang: Efficient execution of structured language model programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng · 2024
Closest in time.
DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Closest in time.
Nanoflow: Towards optimal large language model serving throughput
Kan Zhu, Yilong Zhao, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Yufei Gao, Qinyu Xu, Tian Tang, Zihao Ye, et al · 2024
Closest in time.
Rageval: Scenario specific rag evaluation dataset generation framework, 2024
Kunlun Zhu, Yifan Luo, Dingling Xu, Ruobing Wang, Shi Yu, Shuo Wang, Yukun Yan, Zhenghao Liu, Xu Han, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Reagan: Node-as-agent-reasoning graph agentic network, 2025
Minghao Guo, Xi Zhu, Jingyuan Huang, Kai Mei, and Yongfeng Zhang · 2025
Closest in time.
Mrag-bench: Vision-centric evaluation for retrieval-augmented multimodal models, 2025
Wenbo Hu, Jia-Chen Gu, Zi-Yi Dou, Mohsen Fayyaz, Pan Lu, Kai-Wei Chang, and Nanyun Peng · 2025
Closest in time.
Ket-rag: A cost-efficient multi-granular indexing framework for graph-rag
Yiqian Huang, Shiqi Zhang, and Xiaokui Xiao · 2025
Closest in time.
Rago: Systematic performance optimization for retrieval-augmented generation serving, 2025
Wenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso, Amir Yazdanbakhsh, and Vidushi Dadu · 2025
Closest in time.
Towards agentic rag with deep reasoning: A survey of rag-reasoning systems in llms, 2025
Yangning Li, Weizhi Zhang, Yuyao Yang, Wei-Chieh Huang, Yaozu Wu, Junyu Luo, Yuanchen Bei, Henry Peng Zou, Xiao Luo, Yusheng Zhao, Chunkit Chan, Yankai Chen, Zhongfen Deng, Yinghui Li, Hai-Tao Zheng, Dongyuan Li, Renhe Jiang, Ming Zhang, Yangqiu Song, and Philip S. Yu · 2025
Closest in time.
Telerag: Efficient retrieval-augmented generation inference with lookahead retrieval, 2025
Chien-Yu Lin, Keisuke Kamahori, Yiyu Liu, Xiaoxiang Shi, Madhav Kashyap, Yile Gu, Rulin Shao, Zihao Ye, Kan Zhu, Stephanie Wang, Arvind Krishnamurthy, Rohan Kadekodi, Luis Ceze, and Baris Kasikci · 2025
Closest in time.
Hm-rag: Hierarchical multi-agent multimodal retrieval augmented generation, 2025
Pei Liu, Xin Liu, Ruoyu Yao, Junming Liu, Siyuan Meng, Ding Wang, and Jun Ma · 2025
Closest in time.
Routellm: Learning to route llms with preference data, 2025
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica · 2025
Closest in time.
Agentic retrieval-augmented generation: A survey on agentic rag, 2025
Aditi Singh, Abul Ehtesham, Saket Kumar, and Tala Talaei Khoei · 2025
Closest in time.
In-depth analysis of graph-based rag in a unified framework, 2025
Yingli Zhou, Yaodong Su, Youran Sun, Shu Wang, Taotao Wang, Runyuan He, Yongwei Zhang, Sicong Liang, Xilin Liu, Yuchi Ma, and Yixiang Fang · 2025
Closest in time.