Fetching the paper…
Reading the bibliography…
Retrieval-Augmented Generation (RAG) is often used with Large Language Models (LLMs) to infuse domain knowledge or user-specific information.
INFaaS: A Model-less and Managed Inference Serving System
Francisco Romero, Qian Li, Neeraja J. Yadwadkar, and Christos Kozyrakis. 2020 · 1905
Earlier work this paper cites.
On the jaccard similarity test
GI Ivchenko and SA Honov. 1998 · 1998
Earlier work this paper cites.
Approximate Query Processing: Taming the TeraBytes.. In VLDB , Vol. 10. 645927–672356
Minos N Garofalakis and Phillip B Gibbons. 2001 · 2001
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Looking for a few good metrics: ROUGE and its evaluation. In Ntcir workshop
Chin-Yew Lin and FJ Och. 2004 · 2004
Earlier work this paper cites.
MapReduce: simplified data processing on large clusters
Jeffrey Dean and Sanjay Ghemawat. 2008 · 2008
Earlier work this paper cites.
Applying the roofline model. In 2014 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) . IEEE, 76–85
Georg Ofenbeck, Ruedi Steinmann, Victoria Caparros, Daniele G Spampinato, and Markus Püschel. 2014 · 2014
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
P Rajpurkar. 2016 · 2016
Earlier work this paper cites.
Clipper: A { \{ Low-Latency } \} online prediction serving system. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) . 613–627
Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J Franklin, Joseph E Gonzalez, and Ion Stoica. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani. 2017 · 2017
Earlier work this paper cites.
Live video analytics at scale with approximation and { \{ Delay-Tolerance } \} . In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) . 377–392
Haoyu Zhang, Ganesh Ananthanarayanan, Peter Bodik, Matthai Philipose, Paramvir Bahl, and Michael J Freedman. 2017 · 2017
Earlier work this paper cites.
Real-time change point detection with application to smart home time series data
Samaneh Aminikhanghahi, Tinghui Wang, and Diane J Cook. 2018 · 2018
Earlier work this paper cites.
Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Verdictdb: Universalizing approximate query processing. In Proceedings of the 2018 International Conference on Management of Data . 1461–1476
Yongjoo Park, Barzan Mozafari, Joseph Sorenson, and Junhao Wang. 2018 · 2018
Earlier work this paper cites.
{ \{ VideoChef } \} : Efficient Approximation for Streaming Video Processing Pipelines. In 2018 USENIX Annual Technical Conference (USENIX ATC 18) . 43–56
Ran Xu, Jinkyu Koo, Rakesh Kumar, Peter Bai, Subrata Mitra, Sasa Misailovic, and Saurabh Bagchi. 2018 · 2018
Earlier work this paper cites.
Evaluating question answering evaluation. In Proceedings of the 2nd workshop on machine reading for question answering . 119–124
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Earlier work this paper cites.
Kendall tau sequence distance: Extending Kendall tau from ranks to sequences
Vincent A Cicirello. 2019 · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Earlier work this paper cites.
Nexus: A GPU cluster engine for accelerating DNN-based video analysis. In Proceedings of the 27th ACM Symposium on Operating Systems Principles . 322–337
Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, and Ravi Sundaram. 2019 · 2019
Earlier work this paper cites.
InferLine: ML Prediction Pipeline Provisioning and Management for Tight Latency Objectives
Daniel Crankshaw, Gur-Eyal Sela, Corey Zumar, Xiangxi Mo, Joseph E. Gonzalez, Ion Stoica, and Alexey Tumanov. 2020 · 2020
Earlier work this paper cites.
High-dimensional vector similarity search: from time series to deep network embeddings. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data . 2829–2832
Karima Echihabi. 2020 · 2020
Earlier work this paper cites.
Serving { \{ DNNs } \} like clockwork: Performance predictability from the bottom up. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 443–462
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020 · 2020
Earlier work this paper cites.
Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps. In Proceedings of the 28th International Conference on Computational Linguistics , Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). International Committee on Computational Linguistics, Barcelona, Spain (Online), 6609–6625
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Put an elephant into a fridge: optimizing cache efficiency for in-memory key-value stores
Kefei Wang, Jian Liu, and Feng Chen. 2020 · 2020
Earlier work this paper cites.
NVIDIA A100 Tensor Core GPU: Performance and Innovation
Jack Choquette, Wishwesh Gandhi, Olivier Giroux, Nick Stam, and Ronny Krashinsky. 2021 · 2021
Earlier work this paper cites.
New trends in high-d vector similarity search: al-driven, progressive, and distributed
Karima Echihabi, Kostas Zoumpatianos, and Themis Palpanas. 2021 · 2021
Cited alongside, same era.
ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 692–705
Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W Lee. 2021 · 2021
Cited alongside, same era.
Cape: Encoding relative positions with continuous augmented positional embeddings
Tatiana Likhomanenko, Qiantong Xu, Gabriel Synnaeve, Ronan Collobert, and Alex Rogozhnikov. 2021 · 2021
Cited alongside, same era.
Triton: Open-source GPU programming for neural networks
Philippe Tillet. 2021 · 2021
Cited alongside, same era.
El-attention: Memory efficient lossless attention for generation. In International Conference on Machine Learning . PMLR, 11648–11658
Yu Yan, Jiusheng Chen, Weizhen Qi, Nikhil Bhendawade, Yeyun Gong, Nan Duan, and Ruofei Zhang. 2021 · 2021
Lmsys-chat-1m: A large-scale real-world llm conversation dataset
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P Xing, et al · 2023
Later among the works it cites.
Efficiently Programming Large Language Models using SGLang
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody_Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al · 2023
Later among the works it cites.
Amazon EC2 P4d Instances – AWS
[n. d.] · 2024
Later among the works it cites.
Keyformer: Kv cache reduction through key tokens selection for efficient generative inference
Muhammad Adnan, Akhil Arunkumar, Gaurav Jain, Prashant Nair, Ilya Soloveychik, and Purushotham Kamath. 2024 · 2024
Later among the works it cites.
Approximate Caching for Efficiently Serving { \{ Text-to-Image } \} Diffusion Models. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) . 1173–1189
Shubham Agarwal, Subrata Mitra, Sarthak Chakraborty, Srikrishna Karanam, Koyel Mukherjee, and Shiv Kumar Saini. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022 · 2022
Cited alongside, same era.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022 · 2022
Cited alongside, same era.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022 · 2022
Cited alongside, same era.
Microsecond-scale preemption for concurrent { \{ GPU-accelerated } \} { \{ DNN } \} inferences. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . 539–558
Mingcong Han, Hanze Zhang, Rong Chen, and Haibo Chen. 2022 · 2022
Cited alongside, same era.
A fast post-training pruning framework for transformers
Woosuk Kwon, Sehoon Kim, Michael W Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami. 2022 · 2022
Cited alongside, same era.
xFormers: A modular and hackable Transformer modelling library
Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, Patrick Labatut, Daniel Haziza, Luca Wehrstedt, Jeremy Reizenstein, and Grigory Sizov. 2022 · 2022
Cited alongside, same era.
Dota: detect and omit weak attentions for scalable transformer acceleration. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems . 14–26
Zheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen, Yufei Ding, and Yuan Xie. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Taming { \{ Throughput-Latency } \} Tradeoff in { \{ LLM } \} Inference with { \{ Sarathi-Serve } \} . In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . 117–134
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024 · 2024
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku
AI Anthropic. 2024 · 2024
Later among the works it cites.
The power of noise: Redefining retrieval for rag systems. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 719–729
Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, and Fabrizio Silvestri. 2024 · 2024
Later among the works it cites.
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
Harry Dong, Xinyu Yang, Zhenyu Zhang, Zhangyang Wang, Yuejie Chi, and Beidi Chen. 2024b · 2024
Later among the works it cites.
QAQ: Quality Adaptive Quantization for LLM KV Cache
Shichen Dong, Wen Cheng, Jiayu Qin, and Wei Wang. 2024a · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. 2024 · 2024
Later among the works it cites.
Prompt cache: Modular attention reuse for low-latency inference
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong. 2024 · 2024
Later among the works it cites.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. 2024 · 2024
Later among the works it cites.
RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
Chao Jin, Zili Zhang, Xuanlin Jiang, Fangyue Liu, Xin Liu, Xuanzhe Liu, and Xin Jin. 2024 · 2024
Later among the works it cites.
{ \{ InfiniGen } \} : Efficient generative inference of large language models with dynamic { \{ KV } \} cache management. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . 155–172
Wonbeom Lee, Jungi Lee, Junghwan Seo, and Jaewoong Sim. 2024 · 2024
Later among the works it cites.
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , Kevin Duh, Helena Gomez, and Steven Bethard (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 3799–3821
Shuo Li, Sangdon Park, Insup Lee, and Osbert Bastani. 2024b · 2024
Later among the works it cites.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024c · 2024
Later among the works it cites.
Optimizing llm queries in relational workloads
Shu Liu, Asim Biswal, Audrey Cheng, Xiangxi Mo, Shiyi Cao, Joseph E Gonzalez, Ion Stoica, and Matei Zaharia. 2024a · 2024
Later among the works it cites.
Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time
Zichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang, Victor Xie, Zhaozhuo Xu, Anastasios Kyrillidis, and Anshumali Shrivastava. 2024b · 2024
Later among the works it cites.
RECON: Training-Free Acceleration for Text-to-Image Synthesis with Retrieval of Concept Prompt Trajectories. In European Conference on Computer Vision . Springer, 288–306
Chen-Yi Lu, Shubham Agarwal, Md Mehrab Tanjim, Kanak Mahadik, Anup Rao, Subrata Mitra, Shiv Kumar Saini, Saurabh Bagchi, and Somali Chaterji. 2024 · 2024
Later among the works it cites.
Deepcache: Accelerating diffusion models for free. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15762–15772
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2024 · 2024
Later among the works it cites.
Learning to compress prompts with gist tokens
Jesse Mu, Xiang Li, and Noah Goodman. 2024 · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024 · 2024
Later among the works it cites.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2024
Later among the works it cites.
SlimDB: A space-efficient key-value storage engine for semi-sorted data
Kai Ren, Qing Zheng, Joy Arulraj, and Garth Gibson. 2017 · 2048
Closest in time.