Fetching the paper…
Reading the bibliography…
We show for the first time that large-scale generative pretrained transformer (GPT) family models can be pruned to at least 50% sparsity in one-shot, without any retraining, at minimal loss of accuracy.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1989
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Hassibi, B., Stork, D. G., and Wolff, G. J · 1993
Earlier work this paper cites.
A simple and effective method for removal of hidden units and weights
Hagiwara, M · 1994
Earlier work this paper cites.
The penn treebank: Annotating predicate argument structure
Marcus, M., Kim, G., Marcinkiewicz, M. A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B · 1994
Earlier work this paper cites.
Hypothesising Neural Nets , pp. 81–106
Kingdon, J · 1997
Earlier work this paper cites.
PiQA: An algebra for querying protein data sets
Tata, S. and Patel, J. M · 2003
Earlier work this paper cites.
Iterative thresholding for sparse approximations
Blumensath, T. and Davies, M. E · 2008
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Han, S., Pool, J., Tran, J., and Dally, W. J · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Lsdsem 2017 shared task: The story cloze test
Mostafazadeh, N., Roth, M., Louis, A., Chambers, N., and Allen, J · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Earlier work this paper cites.
A systematic classification of knowledge, reasoning, and context within the ARC dataset
Boratko, M., Padigela, H., Mikkilineni, D., Yuvraj, P., Das, R., McCallum, A., Chang, M., Fokoue-Nkoutche, A., Kapanipathi, P., Mattei, N., et al · 2018
Earlier work this paper cites.
AMC: AutoML for model compression and acceleration on mobile devices
He, Y., Lin, J., Liu, Z., Wang, H., Li, L.-J., and Han, S · 2018
Earlier work this paper cites.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Cited alongside, same era.
Fast sparse convnets
Elsen, E., Dukhan, M., Gale, T., and Simonyan, K · 2020
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Cited alongside, same era.
Inducing and exploiting activation sparsity for fast inference on deep neural networks
Kurtz, M., Kopinsky, J., Gelashvili, R., Matveev, A., Carr, J., Goin, M., Leiserson, W., Moore, S., Nell, B., Shavit, N., and Alistarh, D · 2020
Cited alongside, same era.
Learning N:M fine-grained structured sparse neural networks from scratch
Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., and Li, H · 2021
Later among the works it cites.
The case for 4-bit precision: k-bit inference scaling laws
Dettmers, T. and Zettlemoyer, L · 2022
Later among the works it cites.
LLM.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Later among the works it cites.
EleutherAI LM Evaluation Harness, 2022
EleutherAI · 2022
Later among the works it cites.
SPDY: Accurate pruning with speedup guarantees
Frantar, E. and Alistarh, D · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Up or down? Adaptive rounding for post-training quantization
Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P · 2020
Cited alongside, same era.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A. M · 2020
Cited alongside, same era.
WoodFisher: Efficient second-order approximation for neural network compression
Singh, S. P. and Alistarh, D · 2020
Cited alongside, same era.
M-FAC: Efficient matrix-free approximations of second-order information
Frantar, E., Kurtic, E., and Alistarh, D · 2021
Cited alongside, same era.
Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A · 2021
Cited alongside, same era.
BRECQ: Pushing the limit of post-training quantization by block reconstruction
Li, Y., Gong, R., Tan, X., Yang, Y., Hu, P., Zhang, Q., Yu, F., Wang, W., and Gu, S · 2021
Cited alongside, same era.
Frantar, E., Singh, S. P., and Alistarh, D · 2022
Later among the works it cites.
HuggingFace Perplexity Calculation, 2022
HuggingFace · 2022
Later among the works it cites.
Gmp*: Well-tuned global magnitude pruning can outperform most bert-pruning methods
Kurtic, E. and Alistarh, D · 2022
Later among the works it cites.
The Optimal BERT Surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D · 2022
Later among the works it cites.
A fast post-training pruning framework for transformers
Kwon, W., Kim, S., Mahoney, M. W., Hassoun, J., Keutzer, K., and Gholami, A · 2022
Later among the works it cites.
DeepSparse, 2022
NeuralMagic · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2022
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S · 2022
Later among the works it cites.
ZeroQuant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Later among the works it cites.
OPT: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.