Fetching the paper…
Reading the bibliography…
Long-context models are essential for many applications but face inefficiencies in loading large KV caches during decoding.
Compressive transformers for long-range sequence modelling, 2019
Rae, J. W., Potapenko, A., Jayakumar, S. M., and Lillicrap, T. P · 1911
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L · 2017
Earlier work this paper cites.
The NarrativeQA reading comprehension challenge
Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers
Dasigi, P., Lo, K., Beltagy, I., Cohan, A., Smith, N. A., and Gardner, M · 2021
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Earlier work this paper cites.
How long can open-source llms truly promise on context length?, June 2023
Li, D., Shao, R., Xie, A., Sheng, Y., Zheng, L., Gonzalez, J. E., Stoica, I., Ma, X., and Zhang, H · 2023
Earlier work this paper cites.
Yarn: Efficient context window extension of large language models, 2023
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2023
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Cited alongside, same era.
H 2 o: Heavy-hitter oracle for efficient generative inference of large language models, 2023
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., Wang, Z., and Chen, B · 2023
Cited alongside, same era.
LongBench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., Dong, Y., Tang, J., and Li, J · 2024
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference, 2024
Feng, Y., Lv, J., Cao, Y., Xie, X., and Zhou, S. K · 2024
Later among the works it cites.
Model tells you what to discard: Adaptive kv cache compression for llms, 2024
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., Rodriguez, A., Gregerson, A., Spataru, A., Roziere, B., Biron, B., Tang, B., Chern, B., Caucheteux, C., Nayak, C., Bi, C., Marra, C., et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pyramidkv: Dynamic kv cache compression based on pyramidal information funneling, 2024
Cai, Z., Zhang, Y., Gao, B., Liu, Y., Liu, T., Lu, K., Xiong, W., Dong, Y., Chang, B., Hu, J., and Xiao, W · 2024
Cited alongside, same era.
aws-prototyping/MegaBeam-Mistral-7B-512k, 2024
Chen Wu and Yin Song and Eden Duthie · 2024
Cited alongside, same era.
Clusterkv: Manipulating llm kv cache in semantic space for recallable compression, 2024a
Liu, G., Li, C., Zhao, J., Zhang, C., and Guo, M
Cited in the paper.
Scaling laws of rope-based extrapolation, 2024b
Liu, X., Yan, H., Zhang, S., An, C., Qiu, X., and Lin, D
Cited in the paper.
Later among the works it cites.
Ruler: What’s the real context size of your long-context language models?, 2024
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., and Ginsburg, B · 2024
Later among the works it cites.
Quest: Query-aware sparsity for efficient long-context llm inference, 2024
Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S · 2024
Later among the works it cites.
Flashinfer: Efficient and customizable attention engine for llm inference serving
Ye, Z., Chen, L., Lai, R., Lin, W., Zhang, Y., Wang, S., Chen, T., Kasikci, B., Grover, V., Krishnamurthy, A., and Ceze, L · 2025
Closest in time.