Fetching the paper…
Reading the bibliography…
KV cache stores key and value states from previous tokens to avoid re-computation, yet it demands substantial storage space, especially for long sequences.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Z. Dong, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer · 2019
Earlier work this paper cites.
Haq: Hardware-aware automated quantization with mixed precision
K. Wang, Z. Liu, Y. Lin, J. Lin, and S. Han · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Brecq: Pushing the limit of post-training quantization by block reconstruction
Y. Li, R. Gong, X. Tan, Y. Yang, P. Hu, Q. Zhang, F. Yu, W. Wang, and S. Gu · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Investigating math word problems using pretrained multilingual language models
M. Tan, L. Wang, L. Jiang, and J. Jiang · 2021
Earlier work this paper cites.
Hawq-v3: Dyadic neural network quantization
Z. Yao, Z. Dong, Z. Zheng, A. Gholami, J. Yu, E. Tan, L. Wang, Q. Huang, Y. Wang, M. Mahoney, et al · 2021
Earlier work this paper cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh · 2022
Earlier work this paper cites.
Token dropping for efficient bert pretraining
L. Hou, R. Y. Pang, T. Zhou, Y. Wu, X. Song, X. Song, and D. Zhou · 2022
Earlier work this paper cites.
A systematic evaluation of large language models of code
F. F. Xu, U. Alon, G. Neubig, and V. J. Hellendoorn · 2022
Cited alongside, same era.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Z. Yao, R. Yazdani Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He · 2022
Cited alongside, same era.
Post training mixed precision quantization of neural networks using first-order information
A. Chauhan, U. Tiwari, et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al · 2023
Cited alongside, same era.
FlashAttention-2: Faster attention with better parallelism and work partitioning
T. Dao · 2023
Cited alongside, same era.
Flash-decoding for long-context inference, 2023
Llm-qat: Data-free quantization aware training for large language models
Z. Liu, B. Oguz, C. Zhao, E. Chang, P. Stock, Y. Mehdad, Y. Shi, R. Krishnamoorthi, and V. Chandra · 2023
Later among the works it cites.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
H. Luo, Q. Sun, C. Xu, P. Zhao, J. Lou, C. Tao, X. Geng, Q. Lin, S. Chen, and D. Zhang · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Dao, D. Haziza, F. Massa, and G. Sizov · 2023
Cited alongside, same era.
Can chatgpt pass high school exams on english language comprehension?
J. C. de Winter · 2023
Cited alongside, same era.
Shortcut learning of large language models in natural language understanding
M. Du, F. He, N. Zou, D. Tao, and X. Hu · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou · 2023
Cited alongside, same era.
Ptqd: Accurate post-training quantization for diffusion models
Y. He, L. Liu, J. Liu, W. Wu, H. Zhou, and B. Zhuang · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Cited alongside, same era.
Finequant: Unlocking efficiency with fine-grained weight-only quantization for llms
Y. J. Kim, R. Henry, R. Fahim, and H. H. Awadalla · 2023
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis · 2023
Later among the works it cites.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett, et al · 2023
Later among the works it cites.
Model tells you what to discard: Adaptive kv cache compression for llms
S. Ge, Y. Zhang, L. Liu, M. Zhang, J. Han, and J. Gao · 2024
Closest in time.
Gear: An efficient kv cache compression recipefor near-lossless generative inference of llm
H. Kang, Q. Zhang, S. Kundu, G. Jeong, Z. Liu, T. Krishna, and T. Zhao · 2024
Closest in time.
Common 7b language models already possess strong math capabilities
C. Li, W. Wang, J. Hu, Y. Wei, N. Zheng, H. Hu, Z. Zhang, and H. Peng · 2024
Closest in time.
Qllm: Accurate and efficient low-bitwidth quantization for large language models
J. Liu, R. Gong, X. Wei, Z. Dong, J. Cai, and B. Zhuang · 2024
Closest in time.
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
J. Liu, C. S. Xia, Y. Wang, and L. Zhang · 2024
Closest in time.
Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time
Z. Liu, A. Desai, F. Liao, W. Wang, V. Xie, Z. Xu, A. Kyrillidis, and A. Shrivastava · 2024
Closest in time.
Kivi: A tuning-free asymmetric 2bit quantization for kv cache
Z. Liu, J. Yuan, H. Jin, S. Zhong, Z. Xu, V. Braverman, B. Chen, and X. Hu · 2024
Closest in time.
J. Y. Yang, B. Kim, J. Bae, B. Kwon, G. Park, E. Yang, S. J. Kwon, and D. Lee · 2024
Closest in time.