Fetching the paper…
Reading the bibliography…
The increase in open-source availability of Large Language Models (LLMs) has enabled users to deploy them on more and more resource-constrained edge devices to reduce reliance on network connections and provide more privacy.
2005
Earlier work this paper cites.
IEEE, “IEEE Standard for Standard SystemC Language Reference Manual,” IEEE Std 1666-2011 , 2012
2012
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
O. Zafrir et al. , “Q8BERT: Quantized 8Bit BERT,” in 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing - NeurIPS Edition (EMC2-NIPS) . IEEE Computer Society, Dec. 2019, pp. 36–39
2019
Earlier work this paper cites.
S. Lu et al. , “Hardware Accelerator for Multi-Head Attention and Position-Wise Feed-Forward in the Transformer,” in 2020 IEEE 33rd International System-on-Chip Conference (SOCC) . Las Vegas, NV, USA: IEEE, Sep. 2020, pp. 84–89
2020
Earlier work this paper cites.
“PYNQ Z1 Manual,” https://reference.digilentinc.com/reference/programmable-logic/pynq-z1/reference-manual , 2020
2020
Earlier work this paper cites.
H. Khan et al. , “NPE: An FPGA-based Overlay Processor for Natural Language Processing,” in The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA ’21) . New York, NY, USA: Association for Computing Machinery, Feb. 2021, p. 227
2021
Earlier work this paper cites.
J. Haris et al. , “SECDA: Efficient Hardware/Software Co-Design of FPGA-based DNN Accelerators for Edge Inference,” in 2021 IEEE 33rd International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD) , pp. 33–43. [Online]. Available: https://ieeexplore.ieee.org/document/9651579
2021
Cited alongside, same era.
J. Liu et al. , “Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation,” vol. 36, pp. 21 558–21 572. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2023/hash/43e9d647ccd3e4b7b5baab53f0368686-Abstract-Conference.html
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Kwon et al. , “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , 2023
“GGML,” https://github.com/ggerganov/ggml , 2024
2024
Closest in time.
“llama.cpp,” https://github.com/ggerganov/llama.cpp , 2024
2024
Closest in time.
P. Gibson et al. , “DLAS: A Conceptual Model for Across-Stack Deep Learning Acceleration,” ACM Transactions on Architecture and Code Optimization (TACO) , 2024
2024
Closest in time.
L. J. Wan et al. , “Software/Hardware Co-design for LLM and Its Application for Design Verification,” in 2024 29th Asia and South Pacific Design Automation Conference (ASP-DAC) . Incheon, Korea, Republic of: IEEE, Jan. 2024, pp. 435–441
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Simon Willison, “Alpaca: A New Programming Language for Data Analysis,” {https://simonwillison.net/2023/Mar/13/alpaca/} , 2023, accessed: March 13, 2023
2023
Cited alongside, same era.
Joséphus Cheung, “Guanacodataset (revision 892e57a),” 2023. [Online]. Available: https://huggingface.co/datasets/JosephusCheung/GuanacoDataset
2023
Cited alongside, same era.
P. P. Ray, “ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope,” vol. 3, pp. 121–154. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S266734522300024X
Cited in the paper.
Cited in the paper.
Cited in the paper.
J. Clusmann et al. , “The future landscape of large language models in medicine,” vol. 3, no. 1, pp. 1–8. [Online]. Available: https://www.nature.com/articles/s43856-023-00370-1
Cited in the paper.
J. Haris et al. , “SECDA-TFLite: A toolkit for efficient development of FPGA-based DNN accelerators for edge inference,” vol. 173, pp. 140–151. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0743731522002301
Cited in the paper.