Fetching the paper…
Reading the bibliography…
Various Large Language Models~(LLMs) from the Generative Pretrained Transformer(GPT) family have achieved outstanding performances in a wide range of text generation tasks.
Y. LeCun et al. , “Optimal brain damage,” Neural Information Processing Systems,Neural Information Processing Systems , Jan 1989
1989
Earlier work this paper cites.
B. Hassibi et al. , “Second order derivatives for network pruning: Optimal brain surgeon,” Neural Information Processing Systems,Neural Information Processing Systems , Nov 1992
1992
Earlier work this paper cites.
M. Marcus et al. , “The penn treebank: annotating predicate argument structure,” in Proceedings of the workshop on Human Language Technology - HLT ’94 , Jan 1994
1994
Earlier work this paper cites.
S. Tata et al. , “Piqa: an algebra for querying protein data sets,” in 15th International Conference on Scientific and Statistical Database Management, 2003. , Dec 2003
2003
Earlier work this paper cites.
H. Avron et al. , “Randomized algorithms for estimating the trace of an implicit symmetric positive semi-definite matrix,” Journal of the ACM , p. 1–34, Apr 2011
2011
Earlier work this paper cites.
M. Mahoney, “Randomized algorithms for matrices and data,” Mar 2012
2012
Earlier work this paper cites.
S. Merity et al. , “Pointer sentinel mixture models,” arXiv: Computation and Language,arXiv: Computation and Language , Sep 2016
2016
Earlier work this paper cites.
M. Zhu et al. , “To prune, or not to prune: exploring the efficacy of pruning for model compression,” Oct 2017
2017
Earlier work this paper cites.
M. Boratko et al. , “A systematic classification of knowledge, reasoning, and context within the arc dataset,” Jun 2018
2018
Earlier work this paper cites.
S. Sun et al. , “Patient knowledge distillation for bert model compression,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , Jan 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Z. Dong et al. , “Hawq: Hessian aware quantization of neural networks with mixed-precision,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV) , Oct 2019
2019
Cited alongside, same era.
Z. Dong et al. , “Hawq-v2: Hessian aware trace-weighted quantization of neural networks,” arXiv: Computer Vision and Pattern Recognition,arXiv: Computer Vision and Pattern Recognition , Nov 2019
2019
Cited alongside, same era.
C. Raffel et al. , “Exploring the limits of transfer learning with a unified text-to-text transformer,” arXiv: Learning,arXiv: Learning , Oct 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
I. Hubara et al. , “Accelerated sparse neural training: A provable and efficient method to find n:m transposable masks,” Neural Information Processing Systems,Neural Information Processing Systems , Dec 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
E. Frantar et al. , “Gptq: Accurate post-training quantization for generative pre-trained transformers,” Oct 2022
2022
Later among the works it cites.
Z. Yao et al. , “Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,” Jun 2022
2022
Later among the works it cites.
G. Xiao et al. , “Smoothquant: Accurate and efficient post-training quantization for large language models,” Nov 2022
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Z. Wang et al. , “Structured pruning of large language models,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Jan 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
H. Pan et al. , “Meta-kd: A meta knowledge distillation framework for language model compression across domains,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , Jan 2021
2021
Cited alongside, same era.
H. Bai et al. , “Binarybert: Pushing the limit of bert quantization,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , Jan 2021
2021
Cited alongside, same era.
Z. Liu et al. , “Post-training quantization for vision transformer,” Advances in Neural Information Processing Systems , vol. 34, pp. 28 092–28 103, 2021
2021
Cited alongside, same era.
F. Lagunas et al. , “Block pruning for faster transformers,” arXiv preprint arXiv:2109.04838 , 2021
2021
Cited alongside, same era.
Later among the works it cites.
E. Frantar et al. , “Optimal brain compression: A framework for accurate post-training quantization and pruning,” Aug 2022
2022
Later among the works it cites.
T. Dettmers et al. , “Spqr: A sparse-quantized representation for near-lossless llm weight compression,” Jun 2023
2023
Closest in time.
E. Frantar et al. , “Sparsegpt: Massive language models can be accurately pruned in one-shot,” Jan 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Baichuan-Inc, “Baichuan-7b,” 2023, accessed on Sep 4th, 2023
2023
Closest in time.