Fetching the paper…
Reading the bibliography…
Generative Large Language Models (LLMs) have demonstrated remarkable results for a wide range of tasks.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1990
Earlier work this paper cites.
Latency lags bandwith
Patterson, D. A · 2004
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Balanced csr sparse matrix-vector product on graphics processors
Flegar, G. and Quintana-Ortí, E. S · 2017
Earlier work this paper cites.
Deep neural network compression with single and multiple level quantization, 2018
Xu, Y., Wang, Y., Zhou, A., Lin, W., and Xiong, H · 2018
Earlier work this paper cites.
HAWQ-V2: Hessian Aware trace-Weighted Quantization of neural networks
Dong, Z., Yao, Z., Arfeen, D., Gholami, A., Mahoney, M. W., and Keutzer, K · 2019
Earlier work this paper cites.
Sparse Matrix-Vector Multiplication with CUDA
Evtushenko, G · 2019
Earlier work this paper cites.
Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M · 2019
Earlier work this paper cites.
BinaryBERT: Pushing the limit of BERT quantization
Bai, H., Zhang, W., Hou, L., Shang, L., Jin, J., Jiang, X., Liu, Q., Lyu, M., and King, I · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
ZeroQ: A novel zero shot quantization framework
Cai, Y., Yao, Z., Dong, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Earlier work this paper cites.
Extremely low bit transformer quantization for on-device neural machine translation
Chung, I., Kim, B., Choi, Y., Kwon, S. J., Jeon, Y., Park, B., Kim, S., and Lee, D · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Q-BERT: Hessian based ultra low precision quantization of bert
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Earlier work this paper cites.
GOBO: Quantizing attention-based nlp models for low latency and energy efficient inference
Zadeh, A. H., Edo, I., Awad, O. M., and Moshovos, A · 2020
Earlier work this paper cites.
TernaryBERT: Distillation-aware ultra-low bit bert
Zhang, W., Hou, L., Yin, Y., Shang, L., Chen, X., Jiang, X., and Liu, Q · 2020
Earlier work this paper cites.
Understanding and overcoming the challenges of efficient Transformer quantization
Bondarenko, Y., Nagel, M., and Blankevoort, T · 2021
Earlier work this paper cites.
Scatterbrain: Unifying sparse and low-rank attention
Chen, B., Dao, T., Winsor, E., Song, Z., Rudra, A., and Ré, C · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation, 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., et al · 2021
Cited alongside, same era.
A survey of quantization methods for efficient neural network inference
Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
I-BERT: Integer-only bert quantization
Kim, S., Gholami, A., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Cited alongside, same era.
Bert busters: Outlier dimensions that disrupt transformers
Kovaleva, O., Kulshreshtha, S., Rogers, A., and Rumshisky, A · 2021
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Closest in time.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Closest in time.
Vitality: Unifying low-rank and sparse approximation for vision transformer acceleration with a linear taylor attention
Dass, J., Wu, S., Shi, H., Li, C., Ye, Z., Wang, Z., and Lin, Y · 2023
Closest in time.
SpQR: A sparse-quantized representation for near-lossless LLM weight compression
Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D · 2023
Closest in time.
Output sensitivity-aware detr quantization
Huang, Y., Yang, H., Dong, Z., Gudovskiy, D., Okuno, T., Nakata, Y., Du, Y., Zhang, S., and Keutzer, K · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Post-training sparsity-aware quantization
Shomron, G., Gabbay, F., Kurzum, S., and Weiser, U · 2021
Cited alongside, same era.
Improving neural network quantization without retraining using outlier channel splitting
Zhao, R., Hu, Y., Dotzel, J., De Sa, C., and Zhang, Z · 2021
Cited alongside, same era.
GLAM: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Cited alongside, same era.
GPTQ: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
Mr. BiQ: Post-training non-uniform quantization based on minimizing the reconstruction error
Jeon, Y., Lee, C., Cho, E., and Ro, Y · 2022
Cited alongside, same era.
Non-uniform step size quantization for accurate post-training quantization
Oh, S., Sim, H., Kim, J., and Lee, J · 2022
Cited alongside, same era.
Closest in time.
Full stack optimization of transformer inference: a survey
Kim, S., Hooper, C., Wattanawong, T., Kang, M., Yan, R., Genc, H., Dinh, G., Huang, Q., Keutzer, K., Mahoney, M. W., Shao, S., and Gholami, A · 2023
Closest in time.
Q-diffusion: Quantizing diffusion models
Li, X., Liu, Y., Lian, L., Yang, H., Dong, Z., Kang, D., Zhang, S., and Keutzer, K · 2023
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Closest in time.
NoisyQuant: Noisy bias-enhanced post-training activation quantization for vision transformers
Liu, Y., Yang, H., Dong, Z., Keutzer, K., Du, L., and Zhang, S · 2023
Closest in time.
Omniquant: Omnidirectionally calibrated quantization for large language models
Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K., Gao, P., Qiao, Y., and Luo, P · 2023
Closest in time.
Wei, X., Zhang, Y., Li, Y., Zhang, X., Gong, R., Guo, J., and Liu, X · 2023
Closest in time.
SmoothQuant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Closest in time.
RPTQ: Reorder-based post-training quantization for large language models
Yuan, Z., Niu, L., Liu, J., Liu, W., Wang, X., Shang, Y., Sun, G., Wu, Q., Wu, J., and Wu, B · 2023
Closest in time.
Qd-bev: Quantization-aware view-guided distillation for multi-view 3d object detection
Zhang, Y., Dong, Z., Yang, H., Lu, M., Tseng, C.-C., Guo, Y., Keutzer, K., Du, L., and Zhang, S · 2023
Closest in time.
Quip: 2-bit quantization of large language models with guarantees
Chee, J., Cai, Y., Kuleshov, V., and De Sa, C. M · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2024
Closest in time.
Extreme compression of large language models via additive quantization
Egiazarian, V., Panferov, A., Kuznedelev, D., Frantar, E., Babenko, A., and Alistarh, D · 2024
Closest in time.
Ai and memory wall
Gholami, A., Yao, Z., Kim, S., Hooper, C., Mahoney, M. W., and Keutzer, K · 2024
Closest in time.