Fetching the paper…
Reading the bibliography…
While large language models (LLMs) excel in many domains, their complexity and scale challenge deployment in resource-limited environments.
Y. LeCun, J. Denker, and S. Solla, “Optimal brain damage,” in Advances in Neural Information Processing Systems , D. Touretzky, Ed., vol. 2. Morgan-Kaufmann, 1989. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/1989/file/6c9882bbac1c7093bd25041881277658-Paper.pdf
1989
Earlier work this paper cites.
J. B. Tenenbaum, V. d. Silva, and J. C. Langford, “A global geometric framework for nonlinear dimensionality reduction,” science , vol. 290, no. 5500, pp. 2319–2323, 2000
2000
Earlier work this paper cites.
N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” arXiv preprint physics/0004057 , 2000
2000
Earlier work this paper cites.
I. T. Jolliffe, Principal component analysis for special types of data . Springer, 2002
2002
Earlier work this paper cites.
R. R. Coifman, S. Lafon, A. B. Lee, M. Maggioni, B. Nadler, F. Warner, and S. W. Zucker, “Geometric diffusions as a tool for harmonic analysis and structure definition of data: Diffusion maps,” Proceedings of the national academy of sciences , vol. 102, no. 21, pp. 7426–7431, 2005
2005
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of machine learning research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
T. Hastie, “The elements of statistical learning: data mining, inference, and prediction,” 2009
2009
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning filters for efficient convnets,” 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
L. Deng, G. Li, S. Han, L. Shi, and Y. Xie, “Model compression and hardware acceleration for neural networks: A comprehensive survey,” Proceedings of the IEEE , vol. 108, no. 4, pp. 485–532, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Bisk, R. Zellers, J. Gao, Y. Choi et al. , “Piqa: Reasoning about physical commonsense in natural language,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, no. 05, 2020, pp. 7432–7439
2020
Earlier work this paper cites.
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?” in Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , 2021, pp. 610–623
2021
Cited alongside, same era.
2021
Cited alongside, same era.
P. Ganesh, Y. Chen, X. Lou, M. A. Khan, Y. Yang, H. Sajjad, P. Nakov, D. Chen, and M. Winslett, “Compressing large-scale transformer-based models: A case study on bert,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 1061–1080, 2021
2021
Cited alongside, same era.
A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer, “A survey of quantization methods for efficient neural network inference,” 2021
2021
2023
Later among the works it cites.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning . PMLR, 2023, pp. 38 087–38 099
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer, “Llm.int8(): 8-bit matrix multiplication for transformers at scale,” 2022
2022
Cited alongside, same era.
M. Gupta and P. Agrawal, “Compression of deep learning models for text: A survey,” ACM Transactions on Knowledge Discovery from Data (TKDD) , vol. 16, no. 4, pp. 1–55, 2022
2022
Cited alongside, same era.
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith et al. , “Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time,” in International Conference on Machine Learning . PMLR, 2022, pp. 23 965–23 998
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023
2023
Cited alongside, same era.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al. , “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023
2023
Cited alongside, same era.
X. Zhu, J. Li, Y. Liu, C. Ma, and W. Wang, “A survey on model compression for large language models,” 2023
2023
Cited alongside, same era.
Z. Yang, Y. Zhang, D. Sui, Y. Ju, J. Zhao, and K. Liu, “Explanation guided knowledge distillation for pre-trained language model compression,” ACM Trans. Asian Low-Resour. Lang. Inf. Process. , vol. 23, no. 2, Feb. 2024. [Online]. Available: https://doi.org/10.1145/3639364
2024
Closest in time.
2024
Closest in time.
S. Li, X. Ning, L. Wang, T. Liu, X. Shi, S. Yan, G. Dai, H. Yang, and Y. Wang, “Evaluating quantized large language models,” 2024
2024
Closest in time.
Z. Gong, J. Liu, J. Wang, X. Cai, D. Zhao, and R. Yan, “What makes quantization for large language models hard? an empirical study from the lens of perturbation,” 2024
2024
Closest in time.
2024
Closest in time.
F. Wan, X. Huang, D. Cai, X. Quan, W. Bi, and S. Shi, “Knowledge fusion of large language models,” 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
G. Ortiz-Jimenez, A. Favero, and P. Frossard, “Task arithmetic in the tangent space: Improved editing of pre-trained models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.