Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have significantly advanced the natural language processing paradigm but impose substantial demands on memory and computational resources.
P. Kurup and T. Abbasi, Logic synthesis using Synopsys® . Springer Science & Business Media, 1997
1997
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. H. Oliveira, M. T. Moreira, R. A. Guazzelli, and N. L. Calazans, “Ascend-freepdk45: An open source standard cell library for asynchronous design,” in 2016 IEEE International Conference on Electronics, Circuits and Systems (ICECS) . IEEE, 2016, pp. 652–655
2016
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient processing of deep neural networks: A tutorial and survey,” Proceedings of the IEEE , vol. 105, no. 12, pp. 2295–2329, 2017
2017
Earlier work this paper cites.
H. Kwon, P. Chatarasi, M. Pellauer, A. Parashar, V. Sarkar, and T. Krishna, “Understanding reuse, performance, and hardware cost of dnn dataflow: A data-centric approach,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019, pp. 754–768
2019
Earlier work this paper cites.
M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort, “Up or down? adaptive rounding for post-training quantization,” in International Conference on Machine Learning . PMLR, 2020, pp. 7197–7206
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
D. Wu, J. Li, Z. Pan, Y. Kim, and J. S. Miguel, “ubrain: A unary brain computer interface,” in Proceedings of the 49th Annual International Symposium on Computer Architecture , 2022, pp. 468–481
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Z. Yao, R. Yazdani Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,” Advances in Neural Information Processing Systems , vol. 35, pp. 27 168–27 183, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer, “Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,” Advances in Neural Information Processing Systems , vol. 35, pp. 30 318–30 332, 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quantization for on-device llm compression and acceleration,” Proceedings of Machine Learning and Systems , vol. 6, pp. 87–100, 2024
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Li, X. Ning, K. Hong, T. Liu, L. Wang, X. Li, K. Zhong, G. Dai, H. Yang, and Y. Wang, “Llm-mq: Mixed-precision quantization for efficient llm deployment,” in The Efficient Natural Language and Speech Processing Workshop with NeurIPS , vol. 9, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
C. Guo, J. Tang, W. Hu, J. Leng, C. Zhang, F. Yang, Y. Liu, M. Guo, and Y. Zhu, “Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , 2023, pp. 1–15
2023
Cited alongside, same era.
K. Nassiri and M. Akhloufi, “Transformer models used for text-based question answering systems,” Applied Intelligence , vol. 53, no. 9, pp. 10 602–10 635, 2023
2023
Cited alongside, same era.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 611–626
2023
Cited alongside, same era.
S. Laskaridis, K. Katevas, L. Minto, and H. Haddadi, “Mobile and edge evaluation of large language models,” in Workshop on Efficient Systems for Foundation Models II@ ICML2024
Cited in the paper.
C. Lee, J. Jin, T. Kim, H. Kim, and E. Park, “Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 12, 2024, pp. 13 355–13 364
2024
Later among the works it cites.
2024
Later among the works it cites.
G. Keswani, W. Bisen, H. Padwad, Y. Wankhedkar, S. Pandey, and A. Soni, “Abstractive long text summarization using large language models,” Int. J. Intell. Syst. Appl. Eng , vol. 12, pp. 160–168, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.