Fetching the paper…
Reading the bibliography…
There has been significant interest in "extreme" compression of large language models (LLMs), i.e., to 1-2 bits per parameter, which allows such models to be executed efficiently on resource-constrained devices.
Praktische verfahren der gleichungsauflösung
R. Mises and H. Pollaczek-Geiringer · 1929
Earlier work this paper cites.
The perceptron - a perceiving and recognizing automaton
F. Rosenblatt · 1957
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
30 years of adaptive neural networks: perceptron, madaline, and backpropagation
B. Widrow and M. A. Lehr · 1990
Earlier work this paper cites.
On the convergence of the coordinate descent method for convex differentiable minimization
Z.-Q. Luo and P. Tseng · 1992
Earlier work this paper cites.
Large margin classification using the perceptron algorithm
Y. Freund and R. E. Schapire · 1998
Earlier work this paper cites.
Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation
A. Griewank and A. Walther · 2000
Earlier work this paper cites.
PiQA: An algebra for querying protein data sets
S. Tata and J. M. Patel · 2003
Earlier work this paper cites.
Neural networks for machine learning, coursera (video lectures)., 2012
G. Hinton · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
Additive quantization for extreme vector compression
A. Babenko and V. Lempitsky · 2014
Earlier work this paper cites.
Training deep neural networks with low precision multiplications
M. Courbariaux, Y. Bengio, and J.-P. David · 2014
Earlier work this paper cites.
Stochastic dual ascent for solving linear systems
R. M. Gower and P. Richtárik · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
Parallel coordinate descent methods for big data optimization
P. Richtárik and M. Takáč · 2016
Earlier work this paper cites.
QSGD: Randomized quantization for communication-efficient stochastic gradient descent
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. van den Oord, O. Vinyals, and k. kavukcuoglu · 2017
Earlier work this paper cites.
Efficient inference with TensorRT
H. Vanholder · 2017
Earlier work this paper cites.
The convergence of sparsified gradient methods
D. Alistarh, T. Hoefler, M. Johansson, S. Khirirat, N. Konstantinov, and C. Renggli · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Model compression via distillation and quantization
A. Polino, R. Pascanu, and D. Alistarh · 2018
Earlier work this paper cites.
Latent weights do not exist: Rethinking binarized neural network optimization
K. Helwegen, J. Widdicombe, L. Geiger, Z. Liu, K.-T. Cheng, and R. Nusselder · 2019
Earlier work this paper cites.
D. Kozak, S. Becker, A. Doostan, and L. Tenorio · 2019
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Earlier work this paper cites.
Understanding straight-through estimator in training activation quantized neural nets
P. Yin, J. Lyu, S. Zhang, S. Osher, Y. Qi, and J. Xin · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Cited alongside, same era.
On biased compression for distributed learning
A. Beznosikov, S. Horváth, P. Richtárik, and M. Safaryan · 2020
Cited alongside, same era.
Pytorch distributed: Experiences on accelerating data parallel training, 2020
S. Li, Y. Zhao, R. Varma, O. Salpekar, P. Noordhuis, T. Li, A. Paszke, J. Smith, B. Vaughan, P. Damania, and S. Chintala · 2020
Cited alongside, same era.
Up or down? Adaptive rounding for post-training quantization
M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. Liu · 2020
Cited alongside, same era.
Textbooks are all you need ii: phi-1.5 technical report
Y. Li, S. Bubeck, R. Eldan, A. Del Giorno, S. Gunasekar, and Y. T. Lee · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
J. Lin, J. Tang, H. Tang, S. Yang, X. Dang, and S. Han · 2023
Later among the works it cites.
Specinfer: Accelerating generative llm serving with speculative inference and token tree verification, 2023
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, Z. Wang, R. Y. Y. Wong, Z. Chen, D. Arfeen, R. Abhyankar, and Z. Jia · 2023
Later among the works it cites.
Pb-llm: Partially binarized large language models, 2023
Y. Shang, Z. Yuan, Q. Wu, and Z. Dong · 2023
Later among the works it cites.
Omniquant: Omnidirectionally calibrated quantization for large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zero: memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh · 2020
Cited alongside, same era.
A framework for few-shot language model evaluation, Sept. 2021
L. Gao, J. Tow, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, K. McDonell, N. Muennighoff, J. Phang, L. Reynolds, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou · 2021
Cited alongside, same era.
A survey of quantization methods for efficient neural network inference
A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer · 2021
Cited alongside, same era.
Zero-offload: Democratizing billion-scale model training, 2021
J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He · 2021
Cited alongside, same era.
Winogrande: an adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2021
Cited alongside, same era.
Reintroducing straight-through estimators as principled methods for stochastic binary networks
A. Shekhovtsov and V. Yanush · 2021
Cited alongside, same era.
W. Shao, M. Chen, Z. Zhang, P. Xu, L. Zhao, Z. Li, K. Zhang, P. Gao, Y. Qiao, and P. Luo · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Bitnet: Scaling 1-bit transformers for large language models
H. Wang, S. Ma, L. Dong, S. Huang, H. Wang, L. Ma, F. Yang, R. Wang, Y. Wu, and F. Wei · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone, 2024
M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behl, A. Benhaim, M. Bilenko, J. Bjorck, S. Bubeck, M. Cai, C. C. T. Mendes, W. Chen, V. Chaudhary, P. Chopra, A. D. Giorno, G. de Rosa, M. Dixon, R. Eldan, D. Iter, A. Garg, A. Goswami, S. Gunasekar, E. Haider, J. Hao, R. J. Hewett, J. Huynh, M. Javaheripi, X. Jin, P. Kauffmann, N. Karampatziakis, D. Kim, M. Khademi, L. Kurilenko, J. R. Lee, Y. T. Lee, Y. Li, C. Liang, W. Liu, E. Lin, Z. Lin, P. Madan, A. Mitra, H. Modi, A. Nguyen, B. Norick, B. Patra, D. Perez-Becker, T. Portet, R. Pryzant, H. Qin, M. Radmilac, C. Rosset, S. Roy, O. Ruwase, O. Saarikivi, A. Saied, A. Salim, M. Santacroce, S. Shah, N. Shang, H. Sharma, X. Song, M. Tanaka, X. Wang, R. Ward, G. Wang, P. Witte, M. Wyatt, C. Xu, J. Xu, S. Yadav, F. Yang, Z. Yang, D. Yu, C. Zhang, C. Zhang, J. Zhang, L. L. Zhang, Y. Zhang, Y. Zhang, Y. Zhang, and X. Zhou · 2024
Closest in time.
Fast and optimal weight update for pruned large language models, 2024
V. Boža · 2024
Closest in time.
Db-llm: Accurate dual-binarization for efficient llms, 2024
H. Chen, C. Lv, L. Ding, H. Qin, X. Zhou, Y. Ding, X. Liu, M. Zhang, J. Guo, X. Liu, and D. Tao · 2024
Closest in time.
Extreme compression of large language models via additive quantization
V. Egiazarian, A. Panferov, D. Kuznedelev, E. Frantar, A. Babenko, and D. Alistarh · 2024
Closest in time.
Lq-lora: Low-rank plus quantized matrix decomposition for efficient language model finetuning, 2024
H. Guo, P. Greengard, E. P. Xing, and Y. Kim · 2024
Closest in time.
Billm: Pushing the limit of post-training quantization for llms
W. Huang, Y. Liu, H. Qin, Y. Li, S. Zhang, X. Liu, M. Magno, and X. Qi · 2024
Closest in time.
How good are low-bit quantized llama3 models? an empirical study, 2024
W. Huang, X. Ma, H. Qin, X. Zheng, C. Lv, H. Chen, J. Luo, X. Qi, X. Liu, and M. Magno · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Closest in time.
Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models, 2024
C. Lee, J. Jin, T. Kim, H. Kim, and E. Park · 2024
Closest in time.
The era of 1-bit llms: All large language models are in 1.58 bits, 2024
S. Ma, H. Wang, L. Ma, L. Wang, W. Wang, S. Huang, L. Dong, R. Wang, J. Xue, and F. Wei · 2024
Closest in time.
A simple and effective pruning approach for large language models, 2024
M. Sun, Z. Liu, A. Bair, and J. Z. Kolter · 2024
Closest in time.
Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks
A. Tseng, J. Chee, Q. Sun, V. Kuleshov, and C. De Sa · 2024
Closest in time.
Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks, 2024
A. Tseng, J. Chee, Q. Sun, V. Kuleshov, and C. D. Sa · 2024
Closest in time.
Gptvq: The blessing of dimensionality for llm quantization
M. van Baalen, A. Kuzmin, M. Nagel, P. Couperus, C. Bastoul, E. Mahurin, T. Blankevoort, and P. Whatmough · 2024
Closest in time.
Onebit: Towards extremely low-bit large language models, 2024
Y. Xu, X. Han, Z. Yang, S. Wang, Q. Zhu, Z. Liu, W. Liu, and W. Che · 2024
Closest in time.
Exploring post-training quantization in llms from comprehensive study to low rank compensation
Z. Yao, X. Wu, C. Li, S. Youn, and Y. He · 2024
Closest in time.
Lqer: Low-rank quantization error reconstruction for llms, 2024
C. Zhang, J. Cheng, G. A. Constantinides, and Y. Zhao · 2024
Closest in time.
Tinyllama: An open-source small language model, 2024
P. Zhang, G. Zeng, T. Wang, and W. Lu · 2024
Closest in time.