Fetching the paper…
Reading the bibliography…
Transformer-based Large Language Models (LLMs) have significantly advanced AI capabilities but pose considerable challenges for deployment on edge devices due to high computational demands, memory bandwidth constraints, and energy consumption.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems
2017
Earlier work this paper cites.
Available at https://arxiv.org/abs/1702.03118
S. Elfwing, E. Uchibe, and K. Doya, “Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,” 2017 · 2017
Earlier work this paper cites.
AMD, 2021
AMD Xilinx, Kria K26 SOM Data Sheet · 2021
Earlier work this paper cites.
Available at https://arxiv.org/abs/2108.07258
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, et al · 2022
Earlier work this paper cites.
C. Peng, X. Yang, A. Chen, K. E. Smith, N. PourNejatian, A. B. Costa, C. Martin, M. G. Flores, Y. Zhang, T. Magoc, G. Lipori, D. A. Mitchell, N. S. Ospina, M. M. Ahmed, W. R. Hogan, E. A. Shenkman, Y. Guo, J. Bian, and Y. Wu, “A study of generative large language model for medical research and healthcare,” npj Digital Medicine
2023
Earlier work this paper cites.
Available at https://arxiv.org/abs/2311.07226
F. Zeng, W. Gan, Y. Wang, N. Liu, and P. S. Yu, “Large language models for robotics: A survey,” 2023 · 2023
Earlier work this paper cites.
Available at 10.48550/arXiv.2303.08774
OpenAI, A. Josh, A. Steven, A. Sandhini, A. Lama, A. Ilge, A. Florencia, A. Diogo, A. Janko, A. Sam, A. Shyamal, A. Red, B. Igor, B. Suchir, B. Valerie, B. Paul, B. Haiming, B. Mohammad, B. Jeff, and Z. Barret, “Gpt-4 technical report,” 03 2023 · 2023
Earlier work this paper cites.
Available at https://arxiv.org/abs/2305.10403
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, et al · 2023
Cited alongside, same era.
H. Hua, Y. Li, T. Wang, N. Dong, W. Li, and J. Cao, “Edge computing with artificial intelligence: A machine learning perspective,” ACM Comput. Surv
2023
Cited alongside, same era.
Available at https://arxiv.org/abs/2307.09288
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, et al · 2023
Cited alongside, same era.
Available at https://github.com/karpathy/llama2.c
A. Karpathy, “llama2.c: Inference of llama2 in one file of pure c,” 2023 · 2023
Cited alongside, same era.
R. Yuan, H. Lin, Y. Wang, Z. Tian, S. Wu, T. Shen, et al
2024
Cited alongside, same era.
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, et al
2024
Later among the works it cites.
Available at https://qwenlm.github.io/blog/qwen2.5/
Q. Team, “Qwen2.5: A party of foundation models,” September 2024 · 2024
Later among the works it cites.
H. Chen, J. Zhang, Y. Du, S. Xiang, Z. Yue, N. Zhang, Y. Cai, and Z. Zhang, “Understanding the potential of FPGA-based spatial acceleration for large language model inference,” ACM Transactions on Reconfigurable Technology and Systems
2024
Later among the works it cites.
H. Xu, Y. Li, and S. Ji, “Llamaf: An efficient llama2 architecture accelerator on embedded fpgas,” in 2024 IEEE 10th World Forum on Internet of Things (WF-IoT)
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DeepSeek-AI, “Deepseek-v3 technical report,” 2024 · 2024
Cited alongside, same era.
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quantization for llm compression and acceleration,” in MLSys
2024
Cited alongside, same era.
P. Zhang, G. Zeng, T. Wang, and W. Lu, “Tinyllama: An open-source small language model,” 2024 · 2024
Later among the works it cites.
W. Lan, Z. Tang, M. Liu, Q. Chen, W. Peng, Y. P. Chen, and Y. Pan, “The large language models on biomedical data analysis: A survey,” IEEE Journal of Biomedical and Health Informatics
2025
Closest in time.