Fetching the paper…
Reading the bibliography…
The deployment of large language models (LLMs) is often constrained by memory bandwidth, where the primary bottleneck is the cost of transferring model parameters from the GPU's global memory to its registers.
Multiplication through a single look-up-table (LUT) in CNN inference computation
Shiyu Xu, Qi Wang, Xingbo Wang, Shihang Wang, and Terry Tao Ye. 2021 · 1928
Earlier work this paper cites.
Convolutional neural networks using logarithmic data representation
Daisuke Miyashita, Edward H Lee, and Boris Murmann. 2016 · 2016
Earlier work this paper cites.
Incremental network quantization: Towards lossless CNNs with low-precision weights
Aojun Zhou, Anbang Yao, Yiwen Guo, Lin Xu, and Yurong Chen. 2017 · 2017
Earlier work this paper cites.
LQ-Nets: Learned quantization for highly accurate and compact deep neural networks
Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, and Gang Hua. 2018 · 2018
Earlier work this paper cites.
Quantization networks
Jiwei Yang, Xu Shen, Jun Xing, Xinmei Tian, Houqiang Li, Bing Deng, Jianqiang Huang, and Xian-sheng Hua. 2019 · 2019
Earlier work this paper cites.
Iteratively training look-up tables for network quantization
Fabien Cardinaux, Stefan Uhlich, Kazuki Yoshiyama, Javier Alonso García, Lukas Mauch, Stephen Tiedemann, Thomas Kemp, and Akira Nakamura. 2020 · 2020
Earlier work this paper cites.
Learnable companding quantization for accurate low-bit neural networks
Kohei Yamamoto. 2021 · 2021
Earlier work this paper cites.
LLM.int8(): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
GPTQ: Accurate post-training compression for generative pretrained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022 · 2022
Earlier work this paper cites.
Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, and Dongsoo Lee. 2022 · 2022
Earlier work this paper cites.
Learnable lookup table for neural network quantization
Longguang Wang, Xiaoyu Dong, Yingqian Wang, Li Liu, Wei An, and Yulan Guo. 2022 · 2022
Earlier work this paper cites.
Outlier suppression: Pushing the limit of low-bit transformer language models
Xiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong, Shanghang Zhang, Qi Zhang, Fengwei Yu, and Xianglong Liu. 2022 · 2022
Cited alongside, same era.
SmoothQuant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2022 · 2022
Cited alongside, same era.
Towards end-to-end 4-bit inference on generative large language models
Saleh Ashkboos, Ilia Markov, Elias Frantar, Tingxuan Zhong, Xincheng Wang, Jie Ren, Torsten Hoefler, and Dan Alistarh. 2023 · 2023
Cited alongside, same era.
QLoRA: Efficient finetuning of quantized LLMs
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023 · 2023
Cited alongside, same era.
Efficientqat: Efficient quantization-aware training for large language models
Mengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang, Peng Gao, Kaipeng Zhang, and Ping Luo. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Extreme compression of large language models via additive quantization
Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar, Artem Babenko, and Dan Alistarh. 2024 · 2024
Closest in time.
Marlin: Mixed-precision auto-regressive parallel inference on large language models
Elias Frantar, Roberto L Castro, Jiale Chen, Torsten Hoefler, and Dan Alistarh. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer. 2023 · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
AWQ: Activation-aware weight quantization for LLM compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023 · 2023
Cited alongside, same era.
Stream-K: Work-centric parallel decomposition for dense matrix-matrix multiplication on the GPU
Muhammad Osama, Duane Merrill, Cris Cecka, Michael Garland, and John D Owens. 2023 · 2023
Cited alongside, same era.
OmniQuant: Omnidirectionally calibrated quantization for large language models
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. 2023 · 2023
Cited alongside, same era.
Atom: Low-bit quantization for efficient and accurate LLM serving
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2023 · 2023
Cited alongside, same era.
Quarot: Outlier-free 4-bit inference in rotated LLMs
Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L Croci, Bo Li, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman. 2024 · 2024
Cited alongside, same era.
The era of 1-bit LLMs: All large language models are in 1.58 bits
Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, and Furu Wei. 2024a
Cited in the paper.
Amir Gholami, Zhewei Yao, Sehoon Kim, Coleman Hooper, Michael W Mahoney, and Kurt Keutzer. 2024 · 2024
Closest in time.
LQ-LoRA: Low-rank plus quantized matrix decomposition for efficient language model finetuning
Han Guo, Philip Greengard, Eric P Xing, and Yoon Kim. 2024 · 2024
Closest in time.
How good are low-bit quantized LLaMA3 models? an empirical study
Wei Huang, Xudong Ma, Haotong Qin, Xingyu Zheng, Chengtao Lv, Hong Chen, Jie Luo, Xiaojuan Qi, Xianglong Liu, and Michele Magno. 2024 · 2024
Closest in time.
QServe: W4A8KV4 quantization and system co-design for efficient LLM serving
Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, and Song Han. 2024 · 2024
Closest in time.
Aniruddha Nrusimha, Mayank Mishra, Naigang Wang, Dan Alistarh, Rameswar Panda, and Yoon Kim. 2024 · 2024
Closest in time.
Quip#: Even better LLM quantization with hadamard incoherence and lattice codebooks
Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, and Christopher De Sa. 2024 · 2024
Closest in time.
Ladder: Enabling efficient low-precision deep learning computing through hardware-aware tensor transformation
Lei Wang, Lingxiao Ma, Shijie Cao, Quanlu Zhang, Jilong Xue, Yining Shi, Ningxin Zheng, Ziming Miao, Fan Yang, Ting Cao, Yuqing Yang, and Mao Yang. 2024 · 2024
Closest in time.