Fetching the paper…
Reading the bibliography…
The emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 1910
Earlier work this paper cites.
A generalization of isolated word recognition using vector quantization
Burton, D., Shore, J., and Buck, J · 1983
Earlier work this paper cites.
Vector quantization
Gray, R · 1984
Earlier work this paper cites.
On the statistical analysis of dirty pictures
Besag, J · 1986
Earlier work this paper cites.
PiQA: An algebra for querying protein data sets
Tata, S. and Patel, J. M · 2003
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2009
Earlier work this paper cites.
Approximate nearest neighbor search by residual vector quantization
Chen, Y., Guan, T., and Wang, C · 2010
Earlier work this paper cites.
Product quantization for nearest neighbor search
Jegou, H., Douze, M., and Schmid, C · 2010
Earlier work this paper cites.
Optimized product quantization
Ge, T., He, K., Ke, Q., and Sun, J · 2013
Earlier work this paper cites.
Cartesian k-means
Norouzi, M. and Fleet, D. J · 2013
Earlier work this paper cites.
Additive quantization for extreme vector compression
Babenko, A. and Lempitsky, V · 2014
Earlier work this paper cites.
Composite quantization for approximate nearest neighbor search
Zhang, T., Du, C., and Wang, J · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Quantization based fast inner product search
Guo, R., Kumar, S., Choromanski, K., and Simcha, D · 2016
Earlier work this paper cites.
Revisiting additive quantization
Martinez, J., Clement, J., Hoos, H. H., and Little, J. J · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Competitive quantization for approximate nearest neighbor search
Ozan, E. C., Kiranyaz, S., and Gabbouj, M · 2016
Earlier work this paper cites.
Performance guaranteed network acceleration via high-order residual quantization, 2017
Li, Z., Ni, B., Zhang, W., Yang, X., and Gao, W · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Balanced quantization: An effective and efficient approach to quantized neural networks
Zhou, S.-C., Wang, Y.-Z., Wen, H., He, Q.-Y., and Zou, Y.-H · 2017
Cited alongside, same era.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Cited alongside, same era.
Lsq++: Lower running time and higher recall in multi-codebook quantization
Martinez, J., Zakhmi, S., Hoos, H. H., and Little, J. J · 2018
Cited alongside, same era.
Deep neural network quantization via layer-wise optimization using limited training data
Chen, S., Wang, W., and Pan, S. J · 2019
Cited alongside, same era.
nuQmm: Quantized matmul for efficient inference of large-scale generative language models
Park, G., Park, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2022
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S · 2022
Later among the works it cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Up or down? Adaptive rounding for post-training quantization
Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P · 2020
Cited alongside, same era.
Glu variants improve transformer, 2020
Shazeer, N · 2020
Cited alongside, same era.
Multiplying matrices without multiplying
Blalock, D. and Guttag, J · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Cited alongside, same era.
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Later among the works it cites.
Quip: 2-bit quantization of large language models with guarantees, 2023
Chee, J., Cai, Y., Kuleshov, V., and Sa, C. D · 2023
Later among the works it cites.
Redpajama: an open dataset for training large language models, 2023
Computer, T · 2023
Later among the works it cites.
Are we there yet? product quantization and its hardware acceleration
Fernández-Marqués, J., AbouElhamayed, A. F., Lane, N. D., and Abdelfattah, M. S · 2023
Later among the works it cites.
Qmoe: Practical sub-1-bit compression of trillion-parameter models
Frantar, E. and Alistarh, D · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Later among the works it cites.
Squeezellm: Dense-and-sparse quantization
Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M. W., and Keutzer, K · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Later among the works it cites.
The Falcon family of large language models
TII UAE · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Pv-tuning: Beyond straight-through estimation for extreme llm compression
Malinovskii, V., Mazur, D., Ilin, I., Kuznedelev, D., Burlachenko, K., Yi, K., Alistarh, D., and Richtarik, P · 2024
Closest in time.
Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks, 2024
Tseng, A., Chee, J., Sun, Q., Kuleshov, V., and Sa, C. D · 2024
Closest in time.