Fetching the paper…
Reading the bibliography…
The extremely high computational and storage demands of large language models have excluded most edge devices, which were widely used for efficient machine learning, from being viable options.
Kaiyuan Guo, Lingzhi Sui, Jiantao Qiu, Jincheng Yu, Junbin Wang, Song Yao, Song Han, Yu Wang, and Huazhong Yang, “Angel-eye: A complete design flow for mapping cnn onto embedded fpga,” IEEE transactions on computer-aided design of integrated circuits and systems , vol. 37, no. 1, pp. 35–47, 2017
2017
Earlier work this paper cites.
A Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Stefan Elfwing, Eiji Uchibe, and Kenji Doya, “Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,” Neural networks , vol. 107, pp. 3–11, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Benjamin John Rosser, “Cocotb: a python-based digital logic verification framework,” in Micro-electronics Section seminar. CERN, Geneva, Switzerland , 2018
2018
Earlier work this paper cites.
Biao Zhang and Rico Sennrich, “Root mean square layer normalization,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
Xiang Chen, Jindong Li, and Yong Zhao, “Hardware resource and computational density efficient cnn accelerator design based on fpga,” in 2021 IEEE International Conference on Integrated Circuits, Technologies and Applications (ICTA) . IEEE, 2021, pp. 204–205
2021
Earlier work this paper cites.
Zejian Liu, Gang Li, and Jian Cheng, “Hardware acceleration of fully quantized bert for efficient natural language processing,” in 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE) . IEEE, 2021, pp. 513–516
2021
Earlier work this paper cites.
Zhengang Li, Mengshu Sun, Alec Lu, Haoyu Ma, Geng Yuan, Yanyue Xie, Hao Tang, Yanyu Li, Miriam Leeser, Zhangyang Wang et al. , “Auto-vit-acc: An fpga-aware automatic acceleration framework for vision transformer with mixed-scheme quantization,” in 2022 32nd International Conference on Field-Programmable Logic and Applications (FPL) . IEEE, 2022, pp. 109–116
2022
Cited alongside, same era.
Seongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee, Minsub Kim, Dongsoo Lee, and Joo-Young Kim, “Dfx: A low-latency multi-fpga appliance for accelerating transformer-based text generation,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2022, pp. 616–630
2022
Cited alongside, same era.
Peiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie, Kenneth Liu, Zhenglun Kong, Xin Meng, Zhengang Li, Xue Lin, Zhenman Fang et al. , “Heatvit: Hardware-efficient adaptive token pruning for vision transformers,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2023, pp. 442–455
2023
Cited alongside, same era.
2024
Later among the works it cites.
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neurocomputing , vol. 568, p. 127063, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han, “Awq: Activation-aware weight quantization for on-device llm compression and acceleration,” Proceedings of Machine Learning and Systems , vol. 6, pp. 87–100, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning . PMLR, 2023, pp. 38 087–38 099
2023
Cited alongside, same era.
Jindong Li, Guobin Shen, Dongcheng Zhao, Qian Zhang, and Yi Zeng, “Firefly: A high-throughput hardware accelerator for spiking neural networks with efficient dsp and memory optimization,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. 31, no. 8, pp. 1178–1191, 2023
2023
Cited alongside, same era.
Yi Zeng, Dongcheng Zhao, Feifei Zhao, Guobin Shen, Yiting Dong, Enmeng Lu, Qian Zhang, Yinqian Sun, Qian Liang, Yuxuan Zhao, Zhuoya Zhao, Hongjian Fang, Yuwei Wang, Yang Li, Xin Liu, Chengcheng Du, Qingqun Kong, Zizhe Ruan, and Weida Bi, “BrainCog: A spiking neural network based, brain-inspired cognitive intelligence engine for brain-inspired AI and brain simulation,” Patterns , p. 100789, Jul. 2023. [Online]. Available: https://doi.org/10.1016/j.patter.2023.100789
2023
Cited alongside, same era.
Jindong Li, Guobin Shen, Dongcheng Zhao, Qian Zhang, and Yi Zeng, “Firefly v2: Advancing hardware support for high-performance spiking neural network with a spatiotemporal fpga accelerator,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , 2024
2024
Cited alongside, same era.
Tenglong Li, Jindong Li, Guobin Shen, Dongcheng Zhao, Qian Zhang, and Yi Zeng, “Firefly-s: Exploiting dual-side sparsity for spiking neural networks acceleration with reconfigurable spatial architecture,” IEEE Transactions on Circuits and Systems I: Regular Papers , 2024
2024
Cited alongside, same era.
“Spinalhdl: Scala based hdl.” [Online]. Available: https://github.com/SpinalHDL/SpinalHDL
Cited in the paper.
“Georgi gerganov. ggerganov/llama.cpp: Port of facebook’s llama model in c/c++.” [Online]. Available: https://github.com/ggerganov/
Cited in the paper.
2024
Later among the works it cites.
2024
Later among the works it cites.
Jindong Li, Tenglong Li, Guobin Shen, Dongcheng Zhao, Qian Zhang, and Yi Zeng, “Revealing untapped dsp optimization potentials for fpga-based systolic matrix engines,” in 2024 34th International Conference on Field-Programmable Logic and Applications (FPL) . IEEE, 2024, pp. 197–203
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Nobel Dhar, Bobin Deng, Dan Lo, Xiaofeng Wu, Liang Zhao, and Kun Suo, “An empirical analysis and resource footprint study of deploying large language models on edge devices,” in Proceedings of the 2024 ACM Southeast Conference , 2024, pp. 69–76
2024
Later among the works it cites.