Fetching the paper…
Reading the bibliography…
A critical approach for efficiently deploying computationally demanding large language models (LLMs) is Key-Value (KV) caching.
K. Shoemake, “Animating rotation with quaternion curves,” in Proceedings of the 12th annual conference on Computer graphics and interactive techniques
1985
Earlier work this paper cites.
D. H. Eberly, “Quaternion algebra and calculus,” 2002
2002
Earlier work this paper cites.
M. Roemmele, C. A. Bejan, and A. S. Gordon, “Choice of plausible alternatives: An evaluation of commonsense causal reasoning.,” in AAAI spring symposium: logical formalizations of commonsense reasoning
2011
Earlier work this paper cites.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in NeurIPS
2016
Earlier work this paper cites.
R. Nallapati, B. Zhou, C. Gulcehre, B. Xiang, et al
2016
Earlier work this paper cites.
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal, “Can a suit of armor conduct electricity? a new dataset for open book question answering,” in EMNLP
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson, “Loss surfaces, mode connectivity, and fast ensembling of dnns,” in NeurIPS
2018
Earlier work this paper cites.
A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi, “Mathqa: Towards interpretable math word problem solving with operation-based formalisms,” in NAACL
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Lee, Y. Lee, J. Kim, A. Kosiorek, S. Choi, and Y. W. Teh, “Set transformer: A framework for attention-based permutation-invariant neural networks,” in ICML
2019
Earlier work this paper cites.
S. Reddy, D. Chen, and C. D. Manning, “Coqa: A conversational question answering challenge,” Transactions of the Association for Computational Linguistics
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
W. Zhou, C. Xu, T. Ge, J. McAuley, K. Xu, and F. Wei, “Bert loses patience: Fast and robust inference with early exit,” in NeurIPS
2020
Earlier work this paper cites.
Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi, “Piqa: Reasoning about physical commonsense in natural language,” in AAAI
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort, “Up or down? adaptive rounding for post-training quantization,” in International Conference on Machine Learning
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi, “Winogrande: An adversarial winograd schema challenge at scale,” Communications of the ACM
2021
Earlier work this paper cites.
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. A. Smith, and L. Kong, “Random feature attention,” in ICLR
2021
Earlier work this paper cites.
X. Ma, X. Kong, S. Wang, C. Zhou, J. May, H. Ma, and L. Zettlemoyer, “Luna: Linear unified nested attention,” in NeurIPS
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
L. Gao, J. Tow, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, K. McDonell, N. Muennighoff, J. Phang, L. Reynolds, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou, “A framework for few-shot language model evaluation,” 2021
2021
Cited alongside, same era.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al
2022
Cited alongside, same era.
T. Schuster, A. Fisch, J. Gupta, M. Dehghani, D. Bahri, V. Tran, Y. Tay, and D. Metzler, “Confident adaptive language modeling,” in NeurIPS
2022
Cited alongside, same era.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning
2023
Later among the works it cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” in NeurIPS
2022
Cited alongside, same era.
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer, “Llm.int8(): 8-bit matrix multiplication for transformers at scale,” in NeurIPS
2022
Cited alongside, same era.
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, et al
2022
Cited alongside, same era.
M. S. Matena and C. A. Raffel, “Merging models with fisher-weighted averaging,” in NeurIPS
2022
Cited alongside, same era.
2022
Cited alongside, same era.
OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, and OTHERS, “Gpt-4 technical report,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Liu, H. Li, K. Du, J. Yao, Y. Cheng, Y. Huang, S. Lu, M. Maire, H. Hoffmann, A. Holtzman, et al
2023
Later among the works it cites.
S. Wei, T. Ye, S. Zhang, Y. Tang, and J. Liang, “Joint token pruning and squeezing towards more aggressive compression of vision transformers,” in CVPR
2023
Later among the works it cites.
2023
Later among the works it cites.
Accessed: 2024-05-04
“Introducing meta llama 3: The most capable openly available llm to date.” https://ai.meta.com/blog/meta-llama-3/ , 2024 · 2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett, et al
2024
Closest in time.
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” in ICLR
2024
Closest in time.
S. Ge, Y. Zhang, L. Liu, M. Zhang, J. Han, and J. Gao, “Model tells you what to discard: Adaptive kv cache compression for llms,” ICLR
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Wu and K. Tu, “Layer-condensed kv cache for efficient inference of large language models,” 2024
2024
Closest in time.
D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu, et al
2024
Closest in time.
2024
Closest in time.
J. Liu, R. Gong, X. Wei, Z. Dong, J. Cai, and B. Zhuang, “Qllm: Accurate and efficient low-bitwidth quantization for large language models,” in ICLR
2024
Closest in time.
M. Zhang, C. Shen, Z. Yang, L. Ou, X. Yu, B. Zhuang, et al
2024
Closest in time.
J. Liu, R. Gong, X. Wei, Z. Dong, J. Cai, and B. Zhuang, “Qllm: Accurate and efficient low-bitwidth quantization for large language models,” in ICLR
2024
Closest in time.
2024
Closest in time.
Accessed: 2024-05-04
“Stanford crfm.” https://crfm.stanford.edu/2023/10/12/flashdecoding.html , 2024 · 2024
Closest in time.