Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have pushed the frontier of artificial intelligence but are comprised of hundreds of billions of parameters and operations.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 1909
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models, 2020
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 1910
Earlier work this paper cites.
The mpi message passing interface standard
L. Clarke, I. Glendinning, and R. Hempel · 1994
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
Llm.int8(): 8-bit matrix multiplication for transformers at scale, 2022
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer · 2022
Earlier work this paper cites.
Reducing activation recomputation in large transformer models, 2022
V. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro · 2022
Earlier work this paper cites.
Efficiently scaling transformer inference, 2022
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, A. Levskaya, J. Heek, K. Xiao, S. Agrawal, and J. Dean · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2023
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers, 2023
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh · 2023
Earlier work this paper cites.
Foundation model stack
IBM · 2023
Earlier work this paper cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
J. Lin, J. Tang, H. Tang, S. Yang, X. Dang, and S. Han · 2023
Earlier work this paper cites.
Microscaling data formats for deep learning, 2023
B. D. Rouhani, R. Zhao, A. More, M. Hall, A. Khodamoradi, S. Deng, D. Choudhary, M. Cornea, E. Dellinger, K. Denolf, S. Dusan, V. Elango, M. Golub, A. Heinecke, P. James-Roxby, D. Jani, G. Kolhe, M. Langhammer, A. Li, L. Melnick, M. Mesmakhosroshahi, A. Rodriguez, M. Schulte, R. Shafipour, L. Shao, M. Siu, P. Dubey, P. Micikevicius, M. Naumov, C. Verrilli, R. Wittig, D. Burger, and E. Chung · 2023
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu, 2023
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, D. Y. Fu, Z. Xie, B. Chen, C. Barrett, J. E. Gonzalez, P. Liang, C. Ré, I. Stoica, and C. Zhang · 2023
Cited alongside, same era.
X. Wei, Y. Zhang, Y. Li, X. Zhang, R. Gong, J. Guo, and X. Liu · 2023
Cited alongside, same era.
Gear: An efficient kv cache compression recipe for near-lossless generative inference of llm, 2024
H. Kang, Q. Zhang, S. Kundu, G. Jeong, Z. Liu, T. Krishna, and T. Zhao · 2024
Closest in time.
Large language models: A survey, 2024
S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao · 2024
Closest in time.
microxcaling
Mircosoft · 2024
Closest in time.
Nccl library
Nvidia · 2024
Closest in time.
Blackwell Architecture Overview
NVIDIA · 2024
Closest in time.
Ocp microscaling formats (mx) specification
OCP Specification · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Wu, C. Li, R. Y. Aminabadi, Z. Yao, and Y. He · 2023
Cited alongside, same era.
Rptq: Reorder-based post-training quantization for large language models, 2023
Z. Yuan, L. Niu, J. Liu, W. Liu, X. Wang, Y. Shang, G. Sun, Q. Wu, J. Wu, and B. Wu · 2023
Cited alongside, same era.
Kivi : Plug-and-play 2bit kv cache quantization with streaming asymmetric quantization
Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, V. Braverman, Beidi Chen, and X. Hu · 2023
Cited alongside, same era.
A. Agrawal, J. Chen, Í. Goiri, R. Ramjee, C. Zhang, A. Tumanov, and E. Choukse · 2024
Cited alongside, same era.
Does compressing activations help model parallel training?
S. Bian, D. Li, H. Wang, E. Xing, and S. Venkataraman · 2024
Cited alongside, same era.
The llama 3 herd of models, 2024
A. Dubey et al · 2024
Cited alongside, same era.
Google Cloud Platform
Google · 2024
Cited alongside, same era.
Kvquant: Towards 10 million context length llm inference with kv cache quantization, 2024
C. Hooper, S. Kim, H. Mohammadzadeh, M. W. Mahoney, Y. S. Shao, K. Keutzer, and A. Gholami · 2024
Cited alongside, same era.
torch.compile tutorial
W. Wen · 2024
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models, 2024
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han · 2024
Closest in time.
Atom: Low-bit quantization for efficient and accurate llm serving, 2024
Y. Zhao, C.-Y. Lin, K. Zhu, Z. Ye, L. Chen, S. Zheng, L. Ceze, A. Krishnamurthy, T. Chen, and B. Kasikci · 2024
Closest in time.
On optimizing the communication of model parallelism, 2024
Y. Zhuang, H. Zhao, L. Zheng, Z. Li, E. P. Xing, Q. Ho, J. E. Gonzalez, I. Stoica, and H. Zhang · 2024
Closest in time.