Fetching the paper…
Reading the bibliography…
Deploying large language models (LLMs) on edge platforms is challenged by their high computational and memory demands.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
Y. Qiao, M. Alnemari, and N. Bagherzadeh, “A two-stage efficient 3-d cnn framework for eeg based emotion recognition,” in 2022 IEEE International Conference on Industrial Technology (ICIT) . IEEE, 2022, pp. 1–8
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Ding, Y. Qiao, and N. Bagherzadeh, “Bnn an ideal architecture for acceleration with resistive in memory computation,” IEEE Transactions on Emerging Topics in Computing , vol. 11, no. 2, pp. 281–291, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
R. Sarkar, H. Liang, Z. Fan, Z. Wang, and C. Hao, “Edge-moe: Memory-efficient multi-task vision transformer architecture with task-level sparsity via mixture-of-experts,” in IEEE/ACM International Conference on Computer-Aided Design (ICCAD) , 2023, pp. 1–9
2023
Cited alongside, same era.
2024
Cited alongside, same era.
H. Chen, J. Zhang, Y. Du, S. Xiang, Z. Yue, N. Zhang, Y. Cai, and Z. Zhang, “Understanding the potential of fpga-based spatial acceleration for large language model inference,” ACM Transactions on Reconfigurable Technology and Systems , vol. 18, no. 1, pp. 1–29, 2024
2024
Cited alongside, same era.
2024
2024
Later among the works it cites.
D. Gerlinghoff, B. Choong, R. Goh, W.-F. Wong, and T. Luo, “Table-lookup mac: Scalable processing of quantised neural networks in fpga soft logic,” 04 2024, pp. 235–245
2024
Later among the works it cites.
AMD, “Uram storage bind,” 2024. [Online]. Available: https://docs.amd.com/r/2024.2-English/ug1399-vitis-hls/pragma-HLS-bind_storage
2024
Later among the works it cites.
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Y. Xu, X. Han, Z. Yang, and et al., “Onebit: Towards extremely low-bit large language models,” in Advances in Neural Information Processing Systems (NeurIPS) , 2024
2024
Cited alongside, same era.
D. Du, Y. Zhang, S. Cao, and et al., “Bitdistiller: Unleashing the potential of sub-4-bit llms via self-distillation,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.