Fetching the paper…
Reading the bibliography…
Recently, large language models (LLMs) have achieved huge success in the natural language processing (NLP) field, driving a growing demand to extend their deployment from the cloud to edge devices.
B. Hassibi, D. G. Stork, and G. J. Wolff, “Optimal brain surgeon and general network pruning,” IEEE International Conference on Neural Networks , pp. 293–299 vol.1, 1993
1993
Earlier work this paper cites.
S. Williams, A. Waterman, and D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Commun. ACM , vol. 52, no. 4, p. 65–76, apr 2009
2009
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. He, X. Zhang, and J. Sun, “Channel pruning for accelerating very deep neural networks,” 2017 IEEE International Conference on Computer Vision (ICCV) , pp. 1398–1406, 2017
2017
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019
2019
Earlier work this paper cites.
C. Raffel, N. M. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, pp. 140:1–140:67, 2019
2019
Earlier work this paper cites.
2021
Earlier work this paper cites.
T. J. Ham, Y. Lee, S. H. Seo, S.-U. Kim, H. Choi, S. Jung, and J. W. Lee, “Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks,” 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) , pp. 692–705, 2021
2021
Earlier work this paper cites.
L. Lu, Y. Jin, H. Bi, Z. Luo, P. Li, T. Wang, and Y. Liang, “Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,” MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S.-C. Kao, S. Subramanian, G. Agrawal, A. Yazdanbakhsh, and T. Krishna, “Flat: An optimized dataflow for mitigating attention bottlenecks,” Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2021
2021
Earlier work this paper cites.
Z. Liu, K.-T. Cheng, D. Huang, E. P. Xing, and Z. Shen, “Nonuniform-to-uniform quantization: Towards accurate quantization via generalized straight-through estimation,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 4932–4942, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
R. Y. Aminabadi et al. , “Deepspeed- inference: Enabling efficient inference of transformer models at unprecedented scale,” SC22: International Conference for High Performance Computing, Networking, Storage and Analysis , pp. 1–15, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
S. Hong, S. Moon, J. Kim, S. Lee, M. Kim, D. Lee, and J.-Y. Kim, “Dfx: A low-latency multi-fpga appliance for accelerating transformer-based text generation,” 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) , pp. 616–630, 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
Z. Liu, J. Wang, T. Dao, T. Zhou, B. Yuan, Z. Song, A. Shrivastava, C. Zhang, Y. Tian, C. Ré, and B. Chen, “Deja vu: Contextual sparsity for efficient llms at inference time,” in International Conference on Machine Learning , 2023
2023
Later among the works it cites.
Y. Li, Y. Yu, Q. Zhang, C. Liang, P. He, W. Chen, and T. Zhao, “Losparse: Structured compression of large language models based on low-rank and sparse approximation,” in International Conference on Machine Learning , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Dass, S. Wu, H. Shi, C. Li, Z. Ye, Z. Wang, and Y. Lin, “Vitality: Unifying low-rank and sparse approximation for vision transformer acceleration with a linear taylor attention,” 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , pp. 415–428, 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
O. J. Achiam et al. , “Gpt-4 technical report,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
T. Zhang, F. Ladhak, E. Durmus, P. Liang, K. McKeown, and T. Hashimoto, “Benchmarking large language models for news summarization,” Transactions of the Association for Computational Linguistics , vol. 12, pp. 39–57, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
O. Friha, M. A. Ferrag, B. Kantarci, B. Cakmak, A. Ozgun, and N. Ghoualmi-Zine, “Llm-based edge intelligence: A comprehensive survey on architectures, applications, security and trustworthiness,” IEEE Open Journal of the Communications Society , vol. 5, pp. 5799–5856, 2024
2024
Later among the works it cites.
S. Zeng et al. , “Flightllm: Efficient large language model inference with a complete mapping flow on fpgas,” Proceedings of the 2024 ACM/SIGDA International Symposium on Field Programmable Gate Arrays , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quantization for on-device llm compression and acceleration,” Proceedings of Machine Learning and Systems , vol. 6, pp. 87–100, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
C. Han, Q. Wang, H. Peng, W. Xiong, Y. Chen, H. Ji, and S. Wang, “Lm-infinite: Zero-shot extreme length generalization for large language models,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , 2024, pp. 3991–4008
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
B. Liao and C. Monz, “Apiq: Finetuning of 2-bit quantized large language model,” in Conference on Empirical Methods in Natural Language Processing , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Shi, H. Shao, W. Mao, and Z. Wang, “Trio-vit: Post-training quantization and acceleration for softmax-free efficient vision transformer,” IEEE Transactions on Circuits and Systems I: Regular Papers , vol. 72, pp. 1296–1307, 2024
2024
Later among the works it cites.