Fetching the paper…
Reading the bibliography…
Deploying Large Language Models (LLMs) efficiently on edge devices is often constrained by limited memory capacity and high power consumption.
S. Williams, A. Waterman, and D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Commun. ACM
2009
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” 2015
2015
Earlier work this paper cites.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Binarized neural networks,” NIPS
2016
Earlier work this paper cites.
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “Xnor-net: Imagenet classification using binary convolutional neural networks,” in European conference on computer vision
2016
Earlier work this paper cites.
Y. Dauphin, A. Fan, M. Auli, and D. Grangier, “Language modeling with gated convolutional networks,” 12 2016
2016
Earlier work this paper cites.
A. Vaswani et al
2017
Earlier work this paper cites.
TSMC, “Tsmc’s industry-first and leading 7nm technology enters volume production,” 2018
2018
Earlier work this paper cites.
O. Muller, A. Prost-Boucle, A. Bourge, and F. Pétrot, “Efficient decompression of binary encoded balanced ternary sequences,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems
2019
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root mean square layer normalization,” 12 2019
2019
Earlier work this paper cites.
B. Li, S. Pandey, H. Fang, Y. Lyv, J. Li, J. Chen, M. Xie, L. Wan, H. Liu, and C. Ding, “Ftrans: Energy-efficient acceleration of transformers using fpga,” 2020
2020
Earlier work this paper cites.
Q. Li, X. Zhang, J. Xiong, W.-m. Hwu, and D. Chen, “Efficient methods for mapping neural machine translator on fpgas,” IEEE Transactions on Parallel and Distributed Systems
2020
Earlier work this paper cites.
G. Ilharco, C. Ilharco, I. Turc, T. Dettmers, F. Ferreira, and K. Lee, “High performance natural language processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts
2020
Earlier work this paper cites.
H. Peng, S. Huang, T. Geng, A. Li, W. Jiang, H. Liu, S. Wang, and C. Ding, “Accelerating transformer-based deep learning models on fpgas using column balanced block pruning,” pp. 142–148, 04 2021
2021
Earlier work this paper cites.
P. Qi, Y. Song, H. Peng, S. Huang, Q. Zhuge, and E. Sha, “Accommodating transformer onto fpga: Coupling the balanced model compression and fpga-implementation optimization,” pp. 163–168, 06 2021
2021
Earlier work this paper cites.
F. Li, B. Liu, X. Wang, B. Zhang, and J. Yan, “Ternary weight networks,” 2022
2022
Earlier work this paper cites.
S. Hong, S. Moon, J. Kim, S. Lee, M. Kim, D. Lee, and J.-Y. Kim, “DFX: A low-latency multi-FPGA appliance for accelerating transformer-based text generation,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)
2022
Cited alongside, same era.
S. Hur, S. Na, D. Kwon, J. Kim, J. Kim, A. Boutros, and E. Nurvitadhi, “A fast and flexible fpga-based accelerator for natural language processing neural networks,” ACM Transactions on Architecture and Code Optimization
2022
Cited alongside, same era.
OpenAI, “Chatgpt,” 2023
2023
Cited alongside, same era.
Anthropic, “Claude,” 2023
2023
Cited alongside, same era.
2023
OpenAI, “Gpt-4o mini: Advancing cost-efficient intelligence,” 2024
2024
Later among the works it cites.
A. Gholami, Z. Yao, S. Kim, C. Hooper, M. Mahoney, and K. Keutzer, “Ai and memory wall,” IEEE Micro
2024
Later among the works it cites.
T. Chen, Z. Li, W. Xu, Z. Zhu, D. Li, L. Tian, E. Barsoum, P. Wang, and J. Cheng, “Ternaryllm: Ternarized large language model,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
AMD, “Alveo u280 data center accelerator card data sheet (ds963),” 2023
2023
Cited alongside, same era.
T. Hugo et al., “Llama: Open and efficient foundation language models,” 2023
2023
Cited alongside, same era.
Groq, “12 hours later, groq is running llama 3 instruct 8b & 70b by meta ai on its lpu inference engine,” 2023
2023
Cited alongside, same era.
Google, “Google gemini,” 2024
2024
Cited alongside, same era.
A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al
2024
Cited alongside, same era.
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quantization for on-device llm compression and acceleration,” Proceedings of Machine Learning and Systems
2024
Cited alongside, same era.
2024
Cited alongside, same era.
M. Daliri, Z. Song, and C. Yang, “Unlocking the theory behind scaling 1-bit neural networks,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Wang, H. Zhou, T. Song, S. Mao, S. Ma, H. Wang, Y. Xia, and F. Wei, “1-bit ai infra: Part 1.1, fast and lossless bitnet b1.58 inference on cpus,” 2024
2024
Later among the works it cites.
S. Zeng, J. Liu, G. Dai, X. Yang, T. Fu, H. Wang, W. Ma, H. Sun, S. Li, Z. Huang, et al
2024
Later among the works it cites.
Z. Qin, S. Yang, and Y. Zhong, “Hierarchically gated recurrent neural network for sequence modeling,” Advances in Neural Information Processing Systems
2024
Later among the works it cites.
NVIDIA, “Tensorrt,” 2024
2024
Later among the works it cites.
S. Vosler, “The 200b parameter cruncher macbook pro: Exploring the m4 max llm performance,” 2024
2024
Later among the works it cites.
Cerebras, “Introducing cerebras inference: Ai at instant speed,” Cerebras AI
2024
Later among the works it cites.
Cerebras, “Cerebras product system,” 2024
2024
Later among the works it cites.