Fetching the paper…
Reading the bibliography…
As the size of large language models (LLMs) continues to grow, model compression without sacrificing accuracy has become a crucial challenge for deployment.
Hellaswag: Can a machine really finish your sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 1905
Earlier work this paper cites.
ERNIE: Enhanced language representation with informative entities
Zhang, Z.; Han, X.; Liu, Z.; Jiang, X.; Sun, M.; and Liu, Q. 2019 · 1905
Earlier work this paper cites.
Root Mean Square Layer Normalization
Zhang, B.; and Sennrich, R. 2019 · 1910
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Hassibi, B.; Stork, D. G.; and Wolff, G. J. 1993 · 1993
Earlier work this paper cites.
The Penn treebank: Annotating predicate argument structure
Marcus, M.; Kim, G.; Marcinkiewicz, M. A.; MacIntyre, R.; Bies, A.; Ferguson, M.; Katz, K.; and Schasberger, B. 1994 · 1994
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2016 · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D.; Kruszewski, G.; Lazaridou, A.; Pham, Q. N.; Bernardi, R.; Pezzelle, S.; Baroni, M.; Boleda, G.; and Fernández, R. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T.; Clark, P.; Khot, T.; and Sabharwal, A. 2018 · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y.; Zellers, R.; Gao, J.; Choi, Y.; et al. 2020 · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. 2020 · 2020
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2021 · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation
Gao, L.; Tow, J.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; McDonell, K.; Muennighoff, N.; Phang, J.; Reynolds, L.; Tang, E.; Thite, A.; Wang, B.; Wang, K.; and Zou, A. 2021 · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K.; Bras, R. L.; Bhagavatula, C.; and Choi, Y. 2021 · 2021
Cited alongside, same era.
DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale
Aminabadi, R. Y.; Rajbhandari, S.; Zhang, M.; Awan, A. A.; Li, C.; Li, D.; Zheng, E.; Rasley, J.; Smith, S.; Ruwase, O.; and He, Y. 2022 · 2022
LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation
Li, Y.; Yu, Y.; Zhang, Q.; Liang, C.; He, P.; Chen, W.; and Zhao, T. 2023 · 2023
Closest in time.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Lin, J.; Tang, J.; Tang, H.; Yang, S.; Dang, X.; and Han, S. 2023 · 2023
Closest in time.
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Liu, Z.; Oguz, B.; Zhao, C.; Chang, E.; Stock, P.; Mehdad, Y.; Shi, Y.; Krishnamoorthi, R.; and Chandra, V. 2023 · 2023
Closest in time.
LLM-Pruner: On the Structural Pruning of Large Language Models
Ma, X.; Fang, G.; and Wang, X. 2023 · 2023
Closest in time.
FasterTransformer
NVIDIA. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2022 · 2022
Cited alongside, same era.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Dao, T.; Fu, D. Y.; Ermon, S.; Rudra, A.; and Ré, C. 2022 · 2022
Cited alongside, same era.
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Dettmers, T.; Lewis, M.; Belkada, Y.; and Zettlemoyer, L. 2022 · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E.; Ashkboos, S.; Hoefler, T.; and Alistarh, D. 2022 · 2022
Cited alongside, same era.
The BigScience Corpus: A 1.6 TB Composite Multilingual Dataset
Laurençon, H.; Saulnier, L.; Wang, T.; Akiki, C.; del Moral, A. V.; Le Scao, T.; Von Werra, L.; Mou, C.; Ponferrada, E. G.; Nguyen, H.; et al. 2022 · 2022
Cited alongside, same era.
ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
Yao, Z.; Aminabadi, R. Y.; Zhang, M.; Wu, X.; Li, C.; and He, Y. 2022 · 2022
Cited alongside, same era.
OPT: Open pre-trained transformer language models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022 · 2022
Cited alongside, same era.
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Sheng, Y.; Zheng, L.; Yuan, B.; Li, Z.; Ryabinin, M.; Fu, D. Y.; Xie, Z.; Chen, B.; Barrett, C.; Gonzalez, J. E.; Liang, P.; Ré, C.; Stoica, I.; and Zhang, C. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G.; Lin, J.; Seznec, M.; Wu, H.; Demouth, J.; and Han, S. 2023 · 2023
Closest in time.
Xu, Z.; Liu, Z.; Chen, B.; Tang, Y.; Wang, J.; Zhou, K.; Hu, X.; and Shrivastava, A. 2023 · 2023
Closest in time.
Yao, Z.; Wu, X.; Li, C.; Youn, S.; and He, Y. 2023 · 2023
Closest in time.
RPTQ: Reorder-based Post-training Quantization for Large Language Models
Yuan, Z.; Niu, L.; Liu, J.; Liu, W.; Wang, X.; Shang, Y.; Sun, G.; Wu, Q.; Wu, J.; and Wu, B. 2023 · 2023
Closest in time.
Lifting the Curse of Capacity Gap in Distilling Language Models
Zhang, C.; Yang, Y.; Liu, J.; Wang, J.; Xian, Y.; Wang, B.; and Song, D. 2023 · 2023
Closest in time.