Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly being deployed on edge devices for long-context settings, creating a growing need for fast and efficient long-context inference.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 1911
Earlier work this paper cites.
Roofline: an insightful visual performance model for multicore architectures
Williams, S., Waterman, A., and Patterson, D · 2009
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al · 2020
Earlier work this paper cites.
Transformer acceleration with dynamic sparse attention
Liu, L., Qu, Z., Chen, Z., Ding, Y., and Xie, Y · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Multi-lexsum: Real-world summaries of civil rights lawsuits at multiple granularities
Shen, Z., Lo, K., Yu, L., Dahlberg, N., Schlanger, M., and Downey, D · 2022
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
Earlier work this paper cites.
Flash-decoding for long-context inference: {https://crfm.stanford.edu/2023/10/12/flashdecoding.html} , 2023
Dao, T., Haziza, D., Massa, F., and Sisov, G · 2023
Earlier work this paper cites.
Model tells you what to discard: Adaptive kv cache compression for llms
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Earlier work this paper cites.
Fp8-lm: Training fp8 large language models
Peng, H., Wu, K., Wei, Y., Zhao, G., Yang, Y., Liu, Z., Xiong, Y., Yang, Z., Ni, B., Hu, J., et al · 2023
Earlier work this paper cites.
Omniquant: Omnidirectionally calibrated quantization for large language models
Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K., Gao, P., Qiao, Y., and Luo, P · 2023
Earlier work this paper cites.
Speculative streaming: Fast llm inference without auxiliary models
Bhendawade, N., Belousova, I., Fu, Q., Mason, H., Rastegari, M., and Najibi, M · 2024
Cited alongside, same era.
Reducing transformer key-value cache size with cross-layer attention
Brandon, W., Mishra, M., Nrusimha, A., Panda, R., and Kelly, J. R · 2024
Cited alongside, same era.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T · 2024
Cited alongside, same era.
Quip: 2-bit quantization of large language models with guarantees
Chee, J., Cai, Y., Kuleshov, V., and De Sa, C. M · 2024
Cited alongside, same era.
Dynamic memory compression: Retrofitting llms for accelerated inference
Nawrot, P., Łańcucki, A., Chochowski, M., Tarjan, D., and Ponti, E. M · 2024
Later among the works it cites.
Any-precision llm: Low-cost deployment of multiple, different-sized llms
Park, Y., Hyun, J., Cho, S., Sim, B., and Lee, J. W · 2024
Later among the works it cites.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T · 2024
Later among the works it cites.
Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding
Sun, H., Chen, Z., Yang, X., Tian, Y., and Chen, B · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fishman, M., Chmiel, B., Banner, R., and Soudry, D · 2024
Cited alongside, same era.
Lazyllm: Dynamic token pruning for efficient long context llm inference
Fu, Q., Cho, M., Merth, T., Mehta, S., Rastegari, M., and Najibi, M · 2024
Cited alongside, same era.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A · 2024
Cited alongside, same era.
Kv prediction for improved time to first token, 2024
Horton, M., Cao, Q., Sun, C., Jin, Y., Mehta, S., Rastegari, M., and Nabi, M · 2024
Cited alongside, same era.
Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention
Jiang, H., Li, Y., Zhang, C., Wu, Q., Luo, X., Ahn, S., Han, Z., Abdi, A. H., Li, D., Lin, C.-Y., et al · 2024
Cited alongside, same era.
Gear: An efficient kv cache compression recipefor near-lossless generative inference of llm
Kang, H., Zhang, Q., Kundu, S., Jeong, G., Liu, Z., Krishna, T., and Zhao, T · 2024
Cited alongside, same era.
Speculative decoding with big little decoder
Kim, S., Mangalam, K., Moon, S., Malik, J., Mahoney, M. W., Gholami, A., and Keutzer, K · 2024
Cited alongside, same era.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S · 2024
Cited alongside, same era.
Tan, S., Li, X., Patil, S., Wu, Z., Zhang, T., Keutzer, K., Gonzalez, J. E., and Popa, R. A · 2024
Later among the works it cites.
Quest: Query-aware sparsity for efficient long-context llm inference
Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S · 2024
Later among the works it cites.
Post-training sparse attention with double sparsity
Yang, S., Sheng, Y., Gonzalez, J. E., Stoica, I., and Zheng, L · 2024
Later among the works it cites.
Sirllm: Streaming infinite retentive llm
Yao, Y., Li, Z., and Zhao, H · 2024
Later among the works it cites.
Helmet: How to evaluate long-context language models effectively and thoroughly
Yen, H., Gao, T., Hou, M., Ding, K., Fleischer, D., Izsak, P., Wasserblat, M., and Chen, D · 2024
Later among the works it cites.
Qspec: Speculative decoding with complementary quantization schemes
Zhao, J., Lu, W., Wang, S., Kong, L., and Wu, C · 2024
Later among the works it cites.
Sirius: Contextual sparsity with correction for efficient llms
Zhou, Y., Chen, Z., Xu, Z., Lin, V., and Chen, B · 2024
Later among the works it cites.
Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding
Anonymous · 2025
Closest in time.