Fetching the paper…
Reading the bibliography…
Post-training quantization (PTQ) reduces the memory footprint of LLMs by quantizing their weights to low-precision.
What is the fast fourier transform?
Cochran, W., Cooley, J., Favin, D., Helms, H., Kaenel, R., Lang, W., Maling, G., Nelson, D., Rader, C., and Welch, P · 1967
Earlier work this paper cites.
Unified matrix treatment of the fast walsh-hadamard transform
Fino and Algazi · 1976
Earlier work this paper cites.
Hadamard Matrices and Their Applications
Hedayat, A. and Wallis, W. D · 1978
Earlier work this paper cites.
Multiple stage vector quantization for speech coding
Juang, B.-H. and Gray, A · 1982
Earlier work this paper cites.
Least squares quantization in pcm
Lloyd, S · 1982
Earlier work this paper cites.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
Halko, N., Martinsson, P.-G., and Tropp, J. A · 2011
Earlier work this paper cites.
Fixed-length lossy compression in the finite blocklength regime: Gaussian source
Kostina, V. and Verdú, S · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
The sphere packing problem in dimension 8 8
Viazovska, M · 2017
Earlier work this paper cites.
Up or down? Adaptive rounding for post-training quantization
Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T · 2020
Earlier work this paper cites.
Model preserving compression for neural networks
Chee, J., Renz, M., Damle, A., and Sa, C. D · 2022
Earlier work this paper cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
Overcoming oscillations in quantization-aware training
Nagel, M., Fournarakis, M., Bondarenko, Y., and Blankevoort, T · 2022
Cited alongside, same era.
The falcon series of open language models, 2023
Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Étienne Goffinet, Hesslow, D., Launay, J., Malartic, Q., Mazzotta, D., Noune, B., Pannier, B., and Penedo, G · 2023
Cited alongside, same era.
QuIP: 2-bit quantization of large language models with guarantees
Chee, J., Cai, Y., Kuleshov, V., and Sa, C. D · 2023
Cited alongside, same era.
Redpajama: An open source recipe to reproduce llama training dataset, 2023
Computer, T · 2023
Cited alongside, same era.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Awq: Activation-aware weight quantization for llm compression and acceleration, 2023
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., Gan, C., and Han, S · 2023
Later among the works it cites.
Llm-qat: Data-free quantization aware training for large language models, 2023
Liu, Z., Oguz, B., Zhao, C., Chang, E., Stock, P., Mehdad, Y., Shi, Y., Krishnamoorthi, R., and Chandra, V · 2023
Later among the works it cites.
Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution
Nguyen, E., Poli, M., Faizi, M., Thomas, A., Birch-Sykes, C., Wornow, M., Patel, A., Rabideau, C., Massaroli, S., Bengio, Y., Ermon, S., Baccus, S. A., and Ré, C · 2023
Later among the works it cites.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2023
Later among the works it cites.
Training transformers with 4-bit integers
Xi, H., Li, C., Chen, J., and Zhu, J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dao, T · 2023
Cited alongside, same era.
The case for 4-bit precision: k-bit inference scaling laws
Dettmers, T. and Zettlemoyer, L · 2023
Cited alongside, same era.
Spqr: A sparse-quantized representation for near-lossless llm weight compression, 2023
Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D · 2023
Cited alongside, same era.
OPTQ: Accurate quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
Cited alongside, same era.
Squeezellm: Dense-and-sparse quantization
Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M., and Keutzer, K · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023a
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G
Cited in the paper.
Extreme compression of large language models via additive quantization, 2024
Egiazarian, V., Panferov, A., Kuznedelev, D., Frantar, E., Babenko, A., and Alistarh, D · 2024
Closest in time.
Mixtral of experts, 2024
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.
Code llama: Open foundation models for code, 2024
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Sauvestre, R., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G · 2024
Closest in time.
Omniquant: Omnidirectionally calibrated quantization for large language models
Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K., Gao, P., Qiao, Y., and Luo, P · 2024
Closest in time.
Hadamard Matrices — neilsloane.com
Sloane, N · 2024
Closest in time.