Fetching the paper…
Reading the bibliography…
Scaling the input context length of a large language model (LLM) incurs a significant increase in computation cost and memory footprint to maintain the attention key-value (KV) cache.
Gpgpu processing in cuda architecture
Ghorpade, J., Parande, J., Kulkarni, M., and Bawaskar, A · 2012
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I · 2019
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Llm.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills
Agrawal, A., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., and Ramjee, R · 2023
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Flexgen: High-throughput generative inference of large language models with a single gpu
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Earlier work this paper cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Earlier work this paper cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Earlier work this paper cites.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M., Jacobs, S. A., Awan, A. A., Aneja, J., Awadallah, A., Awadalla, H., Bach, N., Bahree, A., Bakhtiari, A., Behl, H., et al · 2024
Earlier work this paper cites.
L-eval: Instituting standardized evaluation for long context language models
An, C., Gong, S., Zhong, M., Zhao, X., Li, M., Zhang, J., Kong, L., and Qiu, X · 2024
Earlier work this paper cites.
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic · 2024
Earlier work this paper cites.
Longlora: Efficient fine-tuning of long-context large language models
Chen, Y., Qian, S., Tang, H., Lai, X., Liu, Z., Han, S., and Jia, J · 2024
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Earlier work this paper cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Model tells you what to discard: Adaptive kv cache compression for llms
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2024
Cited alongside, same era.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A · 2024
Cited alongside, same era.
Ruler: What’s the real context size of your long-context language models?
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., and Ginsburg, B · 2024
Cited alongside, same era.
Gear: An efficient kv cache compression recipefor near-lossless generative inference of llm
Openai gpt-4o, 2024
OpenAI · 2024
Closest in time.
Anti-haystack, 2024
Pan, W · 2024
Closest in time.
Inference-friendly models with mixattention
Rajput, S., Sheng, Y., Owen, S., and Chiley, V · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Closest in time.
Morehopqa: More than multi-hop reasoning
Schnitzler, J., Ho, X., Huang, J., Boudin, F., Sugawara, S., and Aizawa, A · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kang, H., Zhang, Q., Kundu, S., Jeong, G., Liu, Z., Krishna, T., and Zhao, T · 2024
Cited alongside, same era.
Ktransformers: A flexible framework for experiencing cutting-edge llm inference optimizations, 2024
KVCache.AI · 2024
Cited alongside, same era.
Infinigen: Efficient generative inference of large language models with dynamic kv cache management
Lee, W., Lee, J., Seo, J., and Sim, J · 2024
Cited alongside, same era.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S · 2024
Cited alongside, same era.
Sparser is faster and less is more: Efficient sparse attention for long-range transformers
Lou, C., Jia, Z., Zheng, Z., and Tu, K · 2024
Cited alongside, same era.
Small language models: Survey, measurements, and insights
Lu, Z., Li, X., Cai, D., Yi, R., Liu, F., Zhang, X., Lane, N. D., and Xu, M · 2024
Cited alongside, same era.
Critiprefill: A segment-wise criticality-based approach for prefilling acceleration in llms
Lv, J., Feng, Y., Xie, X., Jia, X., Peng, Q., and Xie, G · 2024
Cited alongside, same era.
Aios: Llm agent operating system
Mei, K., Li, Z., Xu, S., Ye, R., Ge, Y., and Zhang, Y · 2024
Cited alongside, same era.
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T · 2024
Closest in time.
Keep the cost down: A review on methods to optimize llm’s kv-cache consumption
Shi, L., Zhang, H., Yao, Y., Li, Z., and Zhao, H · 2024
Closest in time.
Shadowkv: Kv cache in shadows for high-throughput long-context llm inference
Sun, H., Chang, L.-W., Bao, W., Zheng, S., Zheng, N., Liu, X., Dong, H., Chi, Y., and Chen, B · 2024
Closest in time.
Unlocking longer generation with key-value cache quantization, 2024
Turganbay, R · 2024
Closest in time.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., et al · 2024
Closest in time.
Lm-offload: Performance model-guided generative inference of large language models with parallelism control
Wu, J., Ren, J., Yang, S., Parasyris, K., Georgakoudis, G., Laguna, I., and Li, D · 2024
Closest in time.
Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference
Yang, D., Han, X., Gao, Y., Hu, Y., Zhang, S., and Zhao, H · 2024
Closest in time.
Sirllm: Streaming infinite retentive llm
Yao, Y., Li, Z., and Zhao, H · 2024
Closest in time.
Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches
Yuan, J., Liu, H., Chuang, Y.-N., Li, S., Wang, G., Le, D., Jin, H., Chaudhary, V., Xu, Z., Liu, Z., et al · 2024
Closest in time.
Qjl: 1-bit quantized jl transform for kv cache quantization with zero overhead
Zandieh, A., Daliri, M., and Han, I · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Closest in time.