Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly deployed in large-scale online services, enabling sophisticated applications.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills
Agrawal, A., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., and Ramjee, R · 2023
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Earlier work this paper cites.
Prompt caching, 2023
Anthropic · 2023
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al · 2023
Earlier work this paper cites.
Longlora: Efficient fine-tuning of long-context large language models
Chen, Y., Qian, S., Tang, H., Lai, X., Liu, Z., Han, S., and Jia, J · 2023
Earlier work this paper cites.
Network bandwidth, 2023
Google Cloud · 2023
Earlier work this paper cites.
Llmlingua: Compressing prompts for accelerated inference of large language models
Jiang, H., Wu, Q., Lin, C.-Y., Yang, Y., and Qiu, L · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Earlier work this paper cites.
How long can context length of open-source llms truly promise?
Li, D., Shao, R., Xie, A., Sheng, Y., Zheng, L., Gonzalez, J., Stoica, I., Ma, X., and Zhang, H · 2023
Earlier work this paper cites.
Cachegen: Fast context loading for language model applications
Liu, Y., Li, H., Du, K., Yao, J., Cheng, Y., Huang, Y., Lu, S., Maire, M., Hoffmann, H., Holtzman, A., et al · 2023
Cited alongside, same era.
Skeleton-of-thought: Large language models can do parallel decoding
Ning, X., Lin, Z., Zhou, Z., Wang, Z., Yang, H., and Wang, Y · 2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2023
Cited alongside, same era.
Efficiently programming large language models using sglang
Zheng, L., Yin, L., Xie, Z., Huang, J., Sun, C., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Taming throughput-latency tradeoff in llm inference with sarathi-serve
Longrag: Enhancing retrieval-augmented generation with long-context llms
Jiang, Z., Ma, X., and Chen, W · 2024
Closest in time.
Gear: An efficient kv cache compression recipefor near-lossless generative inference of llm
Kang, H., Zhang, Q., Kundu, S., Jeong, G., Liu, Z., Krishna, T., and Zhao, T · 2024
Closest in time.
Lambda lab gpu cloud specifications, 2024
Lambda · 2024
Closest in time.
Lmcache, 2024
LMCache · 2024
Closest in time.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Zhang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., et al · 2024
Closest in time.
Gpt-4o, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., Tumanov, A., and Ramjee, R · 2024
Cited alongside, same era.
Anthropic model comparison, 2024
Anthropic · 2024
Cited alongside, same era.
Don’t do rag: When cache-augmented generation is all you need for knowledge tasks
Chan, B. J., Chen, C.-T., Cheng, J.-H., and Huang, H.-H · 2024
Cited alongside, same era.
Prompt caching api, 2024
Deepseek · 2024
Cited alongside, same era.
{ \{ Cost-Efficient } \} large language model serving for multi-turn conversations with { \{ CachedAttention } \}
Gao, B., He, Z., Sharma, P., Kang, Q., Jevdjic, D., Deng, J., Yang, X., Yu, Z., and Zuo, P · 2024
Cited alongside, same era.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A · 2024
Cited alongside, same era.
Ragcache: Efficient knowledge caching for retrieval-augmented generation
Jin, C., Zhang, Z., Jiang, X., Liu, F., Liu, X., Liu, X., and Jin, X
Cited in the paper.
Adaptive skeleton graph decoding
Jin, S., Wu, Y., Zheng, H., Zhang, Q., Lentz, M., Mao, Z. M., Prakash, A., Qian, F., and Zhuo, D
Cited in the paper.
openAI · 2024
Closest in time.
Prompt caching, 2024
OpenAI · 2024
Closest in time.
Leave no document behind: Benchmarking long-context llms with extended multi-doc qa
Wang, M., Chen, L., Fu, C., Liao, S., Zhang, X., Wu, B., Yu, H., Xu, N., Zhang, L., Luo, R., et al · 2024
Closest in time.
Cacheblend: Fast large language model serving with cached knowledge fusion
Yao, J., Li, H., Liu, Y., Ray, S., Cheng, Y., Zhang, Q., Du, K., Lu, S., and Jiang, J · 2024
Closest in time.
Chain of agents: Large language models collaborating on long-context tasks
Zhang, Y., Sun, R., Chen, Y., Pfister, T., Zhang, R., and Arik, S. Ö · 2024
Closest in time.