Fetching the paper…
Reading the bibliography…
Post-training quantization (PTQ) is an effective technique for compressing large language models (LLMs).
An algorithm for least-squares estimation of nonlinear parameters
Marquardt, D. W · 1963
Earlier work this paper cites.
Matrix inversion using cholesky decomposition
Krishnamoorthy, A. and Menon, D · 2013
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Dong, Z., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Hawq-v2: Hessian aware trace-weighted quantization of neural networks
Dong, Z., Yao, Z., Arfeen, D., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Up or down? adaptive rounding for post-training quantization
Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Compressing large-scale transformer-based models: A case study on bert
Ganesh, P., Chen, Y., Lou, X., Khan, M. A., Yang, Y., Sajjad, H., Nakov, P., Chen, D., and Winslett, M · 2021
Earlier work this paper cites.
Hawq-v3: Dyadic neural network quantization
Yao, Z., Dong, Z., Zheng, Z., Gholami, A., Yu, J., Tan, E., Wang, L., Huang, Q., Wang, Y., Mahoney, M., et al · 2021
Earlier work this paper cites.
LLM.int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Frantar, E. and Alistarh, D · 2022
Earlier work this paper cites.
GPTQ: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Earlier work this paper cites.
Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R., Zhang, M., Wu, X., Li, C., and He, Y. Z · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
SpQR: A sparse-quantized representation for near-lossless LLM weight compression
Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Qa-lora: Quantization-aware low-rank adaptation of large language models
Xu, Y., Xie, L., Gu, X., Chen, X., Chang, H., Zhang, H., Chen, Z., Zhang, X., and Tian, Q · 2023
Later among the works it cites.
Meta-transformer: A unified framework for multimodal learning
Zhang, Y., Gong, K., Zhang, K., Li, H., Qiao, Y., Ouyang, W., and Yue, X · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
Lq-lora: Low-rank plus quantized matrix decomposition for efficient language model finetuning
Guo, H., Greengard, P., Xing, E. P., and Kim, Y · 2023
Cited alongside, same era.
OWQ: Lessons learned from activation outliers for weight quantization in large language models
Lee, C., Jin, J., Kim, T., Kim, H., and Park, E · 2023
Cited alongside, same era.
Llm-mq: Mixed-precision quantization for efficient llm deployment
Li, S., Ning, X., Hong, K., Liu, T., Wang, L., Li, X., Zhong, K., Dai, G., Yang, H., and Wang, Y · 2023
Cited alongside, same era.
AWQ: Activation-aware weight quantization for LLM compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Cited alongside, same era.
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Liu, Z., Oguz, B., Zhao, C., Chang, E., Stock, P., Mehdad, Y., Shi, Y., Krishnamoorthi, R., and Chandra, V · 2023
Cited alongside, same era.
Ompq: Orthogonal mixed precision quantization
Ma, Y., Jin, T., Zheng, X., Wang, Y., Li, H., Wu, Y., Jiang, G., Zhang, W., and Ji, R · 2023
Cited alongside, same era.
Bibench: Benchmarking and analyzing network binarization
Qin, H., Zhang, M., Ding, Y., Li, A., Cai, Z., Liu, Z., Yu, F., and Liu, X · 2023
Cited alongside, same era.
Zhu, X., Li, J., Liu, Y., Ma, C., and Wang, W · 2023
Later among the works it cites.
A survey on evaluation of large language models
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., et al · 2024
Closest in time.
Quip: 2-bit quantization of large language models with guarantees
Chee, J., Cai, Y., Kuleshov, V., and De Sa, C. M · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2024
Closest in time.
Extreme compression of large language models via additive quantization
Egiazarian, V., Panferov, A., Kuznedelev, D., Frantar, E., Babenko, A., and Alistarh, D · 2024
Closest in time.
APTQ: Attention-aware post-training mixed-precision quantization for large language models
Guan, Z., Huang, H., Su, Y., Huang, H., Wong, N., and Yu, H · 2024
Closest in time.
Apiq: Finetuning of 2-bit quantized large language model
Liao, B. and Monz, C · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Closest in time.
Nrusimha, A., Mishra, M., Wang, N., Alistarh, D., Panda, R., and Kim, Y · 2024
Closest in time.
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
Qin, H., Ma, X., Zheng, X., Li, X., Zhang, Y., Liu, S., Luo, J., Liu, X., and Magno, M · 2024
Closest in time.
Quip#: Even better LLM quantization with hadamard incoherence and lattice codebooks
Tseng, A., Chee, J., Sun, Q., Kuleshov, V., and De Sa, C · 2024
Closest in time.