Fetching the paper…
Reading the bibliography…
Key-value (KV) caching plays an essential role in accelerating decoding for transformer-based autoregressive large language models (LLMs).
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 1904
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma · 2006
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Searching for activation functions, 2017
P. Ramachandran, B. Zoph, and Q. V. Le · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Mixed precision training, 2018
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu · 2018
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le, and R. Salakhutdinov · 2019
Earlier work this paper cites.
GPipe: efficient training of giant neural networks using pipeline parallelism
Y. Huang, Y. Cheng, A. Bapna, O. Firat, M. X. Chen, D. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, and Z. Chen · 2019
Earlier work this paper cites.
A study of bfloat16 for deep learning training, 2019
D. Kalamkar, D. Mudigere, N. Mellempudi, D. Das, K. Banerjee, S. Avancha, D. T. Vooturi, N. Jammalamadaka, J. Huang, H. Yuen, J. Yang, J. Park, A. Heinecke, E. Georganas, S. Srinivasan, A. Kundu, M. Smelyanskiy, B. Kaul, and P. Dubey · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library, 2019
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need, 2019
N. Shazeer · 2019
Earlier work this paper cites.
Neural machine translation with byte-level subwords, 2019
C. Wang, K. Cho, and J. Gu · 2019
Earlier work this paper cites.
Realm: retrieval-augmented language model pre-training
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang · 2020
Earlier work this paper cites.
Transformers are rnns: fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Earlier work this paper cites.
Glu variants improve transformer, 2020
N. Shazeer · 2020
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2020
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model, 2022
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang, M. Pieler, U. S. Prashanth, S. Purohit, L. Reynolds, J. Tow, B. Wang, and S. Weinbach · 2022
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, D. De Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals, S. Osindero, K. Simonyan, J. Rae, E. Elsen, and L. Sifre · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways, 2022
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel · 2022
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett, Z. A. Wang, and B. Chen · 2023
Later among the works it cites.
A survey on model compression for large language models, 2023
X. Zhu, J. Li, Y. Liu, C. Ma, and W. Wang · 2023
Later among the works it cites.
PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation
J. Ansel, E. Yang, H. He, N. Gimelshein, A. Jain, M. Voznesensky, B. Bao, P. Bell, D. Berard, E. Burovski, G. Chauhan, A. Chourdia, W. Constable, A. Desmaison, Z. DeVito, E. Ellison, W. Feng, J. Gong, M. Gschwind, B. Hirsh, S. Huang, K. Kalambarkar, L. Kirsch, M. Lazos, M. Lezcano, Y. Liang, J. Liang, Y. Lu, C. Luk, B. Maher, Y. Pan, C. Puhrsch, M. Reso, M. Saroufim, M. Y. Siraichi, H. Suk, M. Suo, P. Tillet, E. Wang, X. Wang, W. Wen, S. Zhang, X. Zhao, K. Zhou, R. Zou, A. Mathews, G. Chanan, P. Wu, and S. Chintala · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
Memorizing transformers
Y. Wu, M. N. Rabe, D. Hutchins, and C. Szegedy · 2022
Cited alongside, same era.
GQA: Training generalized multi-query transformer models from multi-head checkpoints
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning, 2023
T. Dao · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
A. Gu and T. Dao · 2023
Cited alongside, same era.
Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time, 2023
Z. Liu, A. Desai, F. Liao, W. Wang, V. Xie, Z. Xu, A. Kyrillidis, and A. Shrivastava · 2023
Cited alongside, same era.
Closest in time.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024
DeepSeek-AI · 2024
Closest in time.
Attentionstore: Cost-effective attention reuse across multi-turn conversations in large language model serving, 2024
B. Gao, Z. He, P. Sharma, Q. Kang, D. Jevdjic, J. Deng, X. Yang, Z. Yu, and P. Zuo · 2024
Closest in time.
Model tells you what to discard: Adaptive kv cache compression for llms, 2024
S. Ge, Y. Zhang, L. Liu, M. Zhang, J. Han, and J. Gao · 2024
Closest in time.
Context caching guide
Google · 2024
Closest in time.
Kvquant: Towards 10 million context length llm inference with kv cache quantization, 2024
C. Hooper, S. Kim, H. Mohammadzadeh, M. W. Mahoney, Y. S. Shao, K. Keutzer, and A. Gholami · 2024
Closest in time.
Atlas: few-shot learning with retrieval augmented language models
G. Izacard, P. Lewis, M. Lomeli, L. Hosseini, F. Petroni, T. Schick, J. Dwivedi-Yu, A. Joulin, S. Riedel, and E. Grave · 2024
Closest in time.
Cachegen: Kv cache compression and streaming for fast language model serving, 2024
Y. Liu, H. Li, Y. Cheng, S. Ray, Y. Huang, Q. Zhang, K. Du, J. Yao, S. Lu, G. Ananthanarayanan, M. Maire, H. Hoffmann, A. Holtzman, and J. Jiang · 2024
Closest in time.
Leave no context behind: Efficient infinite context transformers with infini-attention, 2024
T. Munkhdalai, M. Faruqui, and S. Gopal · 2024
Closest in time.
Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence, 2024
B. Peng, D. Goldstein, Q. Anthony, A. Albalak, E. Alcaide, S. Biderman, E. Cheah, X. Du, T. Ferdinan, H. Hou, P. Kazienko, K. K. GV, J. Kocoń, B. Koptyra, S. Krishna, R. M. J. au2, N. Muennighoff, F. Obeid, A. Saito, G. Song, H. Tu, S. Woźniak, R. Zhang, B. Zhao, Q. Zhao, P. Zhou, J. Zhu, and R.-J. Zhu · 2024
Closest in time.
You only cache once: Decoder-decoder architectures for language models, 2024
Y. Sun, L. Dong, Y. Zhu, S. Huang, W. Wang, S. Ma, Q. Zhang, J. Wang, and F. Wei · 2024
Closest in time.
Gated linear attention transformers with hardware-efficient training, 2024
S. Yang, B. Wang, Y. Shen, R. Panda, and Y. Kim · 2024
Closest in time.
Kv cache is 1 bit per channel: Efficient large language model inference with coupled quantization, 2024
T. Zhang, J. Yi, Z. Xu, and A. Shrivastava · 2024
Closest in time.