Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have emerged as powerful tools for natural language processing tasks, revolutionizing the field with their ability to understand and generate human-like text.
A. Stillmaker and B. Baas, “Scaling equations for the accurate prediction of CMOS device performance from 180 nm to 7 nm,” Integration, the VLSI Journal , vol. 58, pp. 74–81, 2017, http://vcl.ece.ucdavis.edu/pubs/2017.02.VLSIintegration.TechScale/
2017
Earlier work this paper cites.
B. Li, S. Pandey, H. Fang, Y. Lyv, J. Li, J. Chen, M. Xie, L. Wan, H. Liu, and C. Ding, “Ftrans: Energy-efficient acceleration of transformers using fpga,” in Proceedings of the ACM/IEEE International Symposium on Low Power Electronics and Design , ser. ISLPED ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 175–180. [Online]. Available: https://doi.org/10.1145/3370748.3406567
2020
Earlier work this paper cites.
S. Lu, M. Wang, S. Liang, J. Lin, and Z. Wang, “Hardware accelerator for multi-head attention and position-wise feed-forward in the transformer,” in 2020 IEEE 33rd International System-on-Chip Conference (SOCC) , 2020, pp. 84–89
2020
Earlier work this paper cites.
T. J. Ham, S. J. Jung, S. Kim, Y. H. Oh, Y. Park, Y. Song, J.-H. Park, S. Lee, K. Park, J. W. Lee, and D.-K. Jeong, “A 3 : Accelerating attention mechanisms in neural networks with approximation,” 2020
2020
Earlier work this paper cites.
H. Guo, L. Peng, J. Zhang, Q. Chen, and T. D. LeCompte, “Att: A fault-tolerant reram accelerator for attention-based neural networks,” in 2020 IEEE 38th International Conference on Computer Design (ICCD) . Los Alamitos, CA, USA: IEEE Computer Society, oct 2020, pp. 213–221. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/ICCD50377.2020.00047
2020
Earlier work this paper cites.
X. Yang, B. Yan, H. Li, and Y. Chen, “Retransformer: Reram-based processing-in-memory architecture for transformer acceleration,” in 2020 IEEE/ACM International Conference On Computer Aided Design (ICCAD) , 2020, pp. 1–9
2020
Earlier work this paper cites.
H. Khan, A. Khan, Z. Khan, L. B. Huang, K. Wang, and L. He, “Npe: An fpga-based overlay processor for natural language processing,” ser. FPGA ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 227. [Online]. Available: https://doi.org/10.1145/3431920.3439477
2021
Earlier work this paper cites.
H. Peng, S. Huang, T. Geng, A. Li, W. Jiang, H. Liu, S. Wang, and C. Ding, “Accelerating transformer-based deep learning models on fpgas using column balanced block pruning,” in 2021 22nd International Symposium on Quality Electronic Design (ISQED) , 2021, pp. 142–148
2021
Earlier work this paper cites.
——, “Accelerating transformer-based deep learning models on fpgas using column balanced block pruning,” in 2021 22nd International Symposium on Quality Electronic Design (ISQED) , 2021, pp. 142–148
2021
Earlier work this paper cites.
J. Fang, Y. Yu, C. Zhao, and J. Zhou, “Turbotransformers: an efficient gpu serving system for transformer models,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , ser. PPoPP ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 389–402. [Online]. Available: https://doi.org/10.1145/3437801.3441578
2021
Earlier work this paper cites.
T. J. Ham, Y. Lee, S. H. Seo, S. Kim, H. Choi, S. J. Jung, and J. W. Lee, “Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) , 2021, pp. 692–705
2021
Earlier work this paper cites.
H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” 2021
2021
Earlier work this paper cites.
L. Lu, Y. Jin, H. Bi, Z. Luo, P. Li, T. Wang, and Y. Liang, “Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,” in MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 977–991. [Online]. Available: https://doi.org/10.1145/3466752.3480125
2021
Earlier work this paper cites.
A. F. Laguna, A. Kazemi, M. Niemier, and X. S. Hu, “In-memory computing based accelerator for transformer networks for long sequences,” in 2021 Design, Automation and Test in Europe Conference and Exhibition (DATE) , 2021, pp. 1839–1844
2021
Earlier work this paper cites.
S. Huang, E. Tang, S. Li, X. Ping, and R. Chen, “Hardware-friendly compression and hardware acceleration for transformer: A survey,” Electronic Research Archive , vol. 30, no. 10, pp. 3755–3785, 2022. [Online]. Available: https://www.aimspress.com/article/doi/10.3934/era.2022192
2022
Earlier work this paper cites.
T. Wang, L. Gong, C. Wang, Y. Yang, Y. Gao, X. Zhou, and H. Chen, “Via: A novel vision-transformer accelerator based on fpga,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 41, no. 11, pp. 4088–4099, 2022
2022
Earlier work this paper cites.
S. Hong, S. Moon, J. Kim, S. Lee, M. Kim, D. Lee, and J.-Y. Kim, “Dfx: A low-latency multi-fpga appliance for accelerating transformer-based text generation,” 2022
2022
Earlier work this paper cites.
C. Fang, S. Guo, W. Wu, J. Lin, Z. Wang, M. K. Hsu, and L. Liu, “An efficient hardware accelerator for sparse transformer neural networks,” in 2022 IEEE International Symposium on Circuits and Systems (ISCAS) , 2022, pp. 2670–2674
2022
Cited alongside, same era.
G. Tzanos, C. Kachris, and D. Soudris, “Hardware acceleration of transformer networks using fpgas,” in 2022 Panhellenic Conference on Electronics and Telecommunications (PACET) , 2022, pp. 1–5
2022
Cited alongside, same era.
J. Choi, H. Li, B. Kim, S. Hwang, and J. H. Ahn, “Accelerating transformer networks through recomposing softmax layers,” in 2022 IEEE International Symposium on Workload Characterization (IISWC) , 2022, pp. 92–103
2022
Cited alongside, same era.
——, “Accelerating transformer networks through recomposing softmax layers,” in 2022 IEEE International Symposium on Workload Characterization (IISWC) , 2022, pp. 92–103
2022
Cited alongside, same era.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , ser. SOSP ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 611–626. [Online]. Available: https://doi.org/10.1145/3600006.3613165
2023
Later among the works it cites.
S. Tuli and N. K. Jha, “Acceltran: A sparsity-aware accelerator for dynamic inference with transformers,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 42, no. 11, pp. 4038–4051, 2023
2023
Later among the works it cites.
Z. Zhou, J. Liu, Z. Gu, and G. Sun, “Energon: Toward efficient acceleration of transformers using dynamic sparse attention,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 42, no. 1, pp. 136–149, 2023
2023
Later among the works it cites.
S. Sridharan, J. R. Stevens, K. Roy, and A. Raghunathan, “X-former: In-memory acceleration of transformers,” 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Wang, Y. Wei, Y. Xiong, G. Huang, X. Qian, Y. Ding, M. Wang, and L. Li, “Lightseq2: Accelerated training for transformer-based models on gpus,” in SC22: International Conference for High Performance Computing, Networking, Storage and Analysis , 2022, pp. 1–14
2022
Cited alongside, same era.
G. Shen, J. Zhao, Q. Chen, J. Leng, C. Li, and M. Guo, “Salo: an efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequences,” in Proceedings of the 59th ACM IEEE Design Automation Conference , ser. DAC ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 571–576. [Online]. Available: https://doi.org/10.1145/3489517.3530504
2022
Cited alongside, same era.
T. Yang, D. Li, Z. Song, Y. Zhao, F. Liu, Z. Wang, Z. He, and L. Jiang, “Dtqatten: Leveraging dynamic token-based quantization for efficient attention architecture,” in 2022 Design, Automation and Test in Europe Conference and Exhibition (DATE) , 2022, pp. 700–705
2022
Cited alongside, same era.
M. Zhou, W. Xu, J. Kang, and T. Rosing, “Transpim: A memory-based acceleration via software-hardware co-design for transformer,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , 2022, pp. 1071–1085
2022
Cited alongside, same era.
J. Zhong, Z. Liu, and X. Chen, “Transformer-based models and hardware acceleration analysis in autonomous driving: A survey,” 2023
2023
Cited alongside, same era.
M. Emani, S. Foreman, V. Sastry, Z. Xie, S. Raskar, W. Arnold, R. Thakur, V. Vishwanath, and M. E. Papka, “A comprehensive performance study of large language models on novel ai accelerators,” 2023
2023
Cited alongside, same era.
Y. Bai, H. Zhou, K. Zhao, J. Chen, J. Yu, and K. Wang, “Transformer-opu: An fpga-based overlay processor for transformer networks,” in 2023 IEEE 31st Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) , 2023, pp. 221–221
2023
Cited alongside, same era.
S. Hur, S. Na, D. Kwon, J. Kim, A. Boutros, E. Nurvitadhi, and J. Kim, “A fast and flexible fpga-based accelerator for natural language processing neural networks,” ACM Trans. Archit. Code Optim. , vol. 20, no. 1, feb 2023. [Online]. Available: https://doi.org/10.1145/3564606
2023
Cited alongside, same era.
2023
Later among the works it cites.
F. Tu, Z. Wu, Y. Wang, L. Liang, L. Liu, Y. Ding, L. Liu, S. Wei, Y. Xie, and S. Yin, “Trancim: Full-digital bitline-transpose cim-based sparse transformer accelerator with pipeline/parallel reconfigurable modes,” IEEE Journal of Solid-State Circuits , vol. 58, no. 6, pp. 1798–1809, 2023
2023
Later among the works it cites.
W. Li, M. Manley, J. Read, A. Kaul, M. S. Bakir, and S. Yu, “H3datten: Heterogeneous 3-d integrated hybrid analog and digital compute-in-memory accelerator for vision transformer self-attention,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. 31, no. 10, pp. 1592–1602, 2023
2023
Later among the works it cites.
I. Okubo, K. Sugiura, and H. Matsutani, “A cost-efficient fpga implementation of tiny transformer model using neural ode,” 2024
2024
Closest in time.
Y. Ji, C. Fang, and Z. Wang, “Beta: Binarized energy-efficient transformer accelerator at the edge,” 2024
2024
Closest in time.
K. Marino, P. Zhang, and V. Prasanna, “Me-vit: A single-load memory-efficient fpga accelerator for vision transformers,” 2024
2024
Closest in time.
D. Danopoulos, G. Zervakis, D. Soudris, and J. Henkel, “Transaxx: Efficient transformers with approximate computing,” 2024
2024
Closest in time.
I. Okubo, K. Sugiura, and H. Matsutani, “A cost-efficient fpga implementation of tiny transformer model using neural ode,” 2024
2024
Closest in time.
J. Zhuang, Z. Yang, S. Ji, H. Huang, A. K. Jones, J. Hu, Y. Shi, and P. Zhou, “Ssr: Spatial sequential hybrid architecture for latency throughput tradeoff in transformer acceleration,” in Proceedings of the 2024 ACM/SIGDA International Symposium on Field Programmable Gate Arrays , ser. FPGA ’24. ACM, Apr. 2024. [Online]. Available: http://dx.doi.org/10.1145/3626202.3637569
2024
Closest in time.
Y. Zhao, D. Wu, and J. Wang, “Alisa: Accelerating large language model inference via sparsity-aware kv caching,” 2024
2024
Closest in time.
Y. Luo and S. Yu, “H3d-transformer: A heterogeneous 3d (h3d) computing platform for transformer model acceleration on edge devices,” ACM Trans. Des. Autom. Electron. Syst. , vol. 29, no. 3, apr 2024. [Online]. Available: https://doi.org/10.1145/3649219
2024
Closest in time.
J. Zhao, P. Zeng, G. Shen, Q. Chen, and M. Guo, “Hardware-software co-design enabling static and dynamic sparse attention mechanisms,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , pp. 1–1, 2024
2024
Closest in time.
Y. Pan, M. Zhou, C. Lee, Z. Li, R. Kushwah, V. Narayanan, and T. Rosing, “Primate: Processing in memory acceleration for dynamic token-pruning transformers,” in 2024 29th Asia and South Pacific Design Automation Conference (ASP-DAC) , 2024, pp. 557–563
2024
Closest in time.
S. Liu, C. Mu, H. Jiang, Y. Wang, J. Zhang, F. Lin, K. Zhou, Q. Liu, and C. Chen, “Hardsea: Hybrid analog-reram clustering and digital-sram in-memory computing accelerator for dynamic sparse self-attention in transformer,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. 32, no. 2, pp. 269–282, 2024
2024
Closest in time.