Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have emerged as powerful tools for natural language processing tasks, revolutionizing the field with their ability to understand and generate human-like text.
CACTI 3.0: An Integrated Cache Timing, Power, and Area Model
Premkishore Shivakumar and Norman Jouppi. 2001 · 2001
Earlier work this paper cites.
A 3 : Accelerating Attention Mechanisms in Neural Networks with Approximation
Tae Jun Ham, Sung Jun Jung, Seonghak Kim, Young H. Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W. Lee, and Deog-Kyoon Jeong. 2020 · 2002
Earlier work this paper cites.
The CUBLAS and CULA based GPU acceleration of adaptive finite element framework for bioluminescence tomography
Bo Zhang, Xiang Yang, Fei Yang, Xin Yang, Chenghu Qin, Dong Han, Xibo Ma, Kai Liu, and Jie Tian. 2010 · 2010
Earlier work this paper cites.
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
Hanrui Wang, Zhekai Zhang, and Song Han. 2021 · 2012
Earlier work this paper cites.
MnnFast: A Fast and Scalable System Architecture for Memory-Augmented Neural Networks. In 2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA) . 250–263
Hanhwi Jang, Joonsung Kim, Jae-Eon Jo, Jaewon Lee, and Jangwoo Kim. 2019 · 2019
Earlier work this paper cites.
ATT: A Fault-Tolerant ReRAM Accelerator for Attention-based Neural Networks. In 2020 IEEE 38th International Conference on Computer Design (ICCD) . IEEE Computer Society, Los Alamitos, CA, USA, 213–221
H. Guo, L. Peng, J. Zhang, Q. Chen, and T. D. LeCompte. 2020 · 2020
Earlier work this paper cites.
FTRANS: Energy-Efficient Acceleration of Transformers Using FPGA. In Proceedings of the ACM/IEEE International Symposium on Low Power Electronics and Design (Boston, Massachusetts) (ISLPED ’20) . Association for Computing Machinery, New York, NY, USA, 175–180
Bingbing Li, Santosh Pandey, Haowen Fang, Yanjun Lyv, Ji Li, Jieyang Chen, Mimi Xie, Lipeng Wan, Hang Liu, and Caiwen Ding. 2020 · 2020
Earlier work this paper cites.
Hardware Accelerator for Multi-Head Attention and Position-Wise Feed-Forward in the Transformer. In 2020 IEEE 33rd International System-on-Chip Conference (SOCC) . 84–89
Siyuan Lu, Meiqi Wang, Shuang Liang, Jun Lin, and Zhongfeng Wang. 2020 · 2020
Earlier work this paper cites.
ReTransformer: ReRAM-based Processing-in-Memory Architecture for Transformer Acceleration. In 2020 IEEE/ACM International Conference On Computer Aided Design (ICCAD) . 1–9
Xiaoxuan Yang, Bonan Yan, Hai Li, and Yiran Chen. 2020 · 2020
Earlier work this paper cites.
ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural Networks. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) . 692–705
Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W. Lee. 2021 · 2021
Earlier work this paper cites.
NPE: An FPGA-Based Overlay Processor for Natural Language Processing (FPGA ’21) . Association for Computing Machinery, New York, NY, USA, 227
Hamza Khan, Asma Khan, Zainab Khan, Lun Bin Huang, Kun Wang, and Lei He. 2021 · 2021
Earlier work this paper cites.
In-Memory Computing based Accelerator for Transformer Networks for Long Sequences. In 2021 Design, Automation and Test in Europe Conference and Exhibition (DATE) . 1839–1844
Ann Franchesca Laguna, Arman Kazemi, Michael Niemier, and X. Sharon Hu. 2021 · 2021
Cited alongside, same era.
Sanger: A Co-Design Framework for Enabling Sparse Attention Using Reconfigurable Architecture. In MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture (Virtual Event, Greece) (MICRO ’21) . Association for Computing Machinery, New York, NY, USA, 977–991
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang. 2021 · 2021
Cited alongside, same era.
Accelerating Transformer-based Deep Learning Models on FPGAs using Column Balanced Block Pruning. In 2021 22nd International Symposium on Quality Electronic Design (ISQED) . 142–148
Hongwu Peng, Shaoyi Huang, Tong Geng, Ang Li, Weiwen Jiang, Hang Liu, Shusen Wang, and Caiwen Ding. 2021 · 2021
Cited alongside, same era.
Accelerating Transformer Networks through Recomposing Softmax Layers. In 2022 IEEE International Symposium on Workload Characterization (IISWC) . 92–103
Accelerating Large Language Model Decoding with Speculative Sampling
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023 · 2023
Later among the works it cites.
A Comprehensive Performance Study of Large Language Models on Novel AI Accelerators
Murali Emani, Sam Foreman, Varuni Sastry, Zhen Xie, Siddhisanket Raskar, William Arnold, Rajeev Thakur, Venkatram Vishwanath, and Michael E. Papka. 2023 · 2023
Later among the works it cites.
Simplifying Transformer Blocks
Bobby He and Thomas Hofmann. 2023 · 2023
Later among the works it cites.
A Fast and Flexible FPGA-Based Accelerator for Natural Language Processing Neural Networks
Suyeon Hur, Seongmin Na, Dongup Kwon, Joonsung Kim, Andrew Boutros, Eriko Nurvitadhi, and Jangwoo Kim. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jaewan Choi, Hailong Li, Byeongho Kim, Seunghwan Hwang, and Jung Ho Ahn. 2022 · 2022
Cited alongside, same era.
DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation
Seongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee, Minsub Kim, Dongsoo Lee, and Joo-Young Kim. 2022 · 2022
Cited alongside, same era.
Hardware-friendly compression and hardware acceleration for transformer: A survey
Shizhen Huang, Enhao Tang, Shun Li, Xiangzhan Ping, and Ruiqi Chen. 2022 · 2022
Cited alongside, same era.
Hardware Acceleration of Transformer Networks using FPGAs. In 2022 Panhellenic Conference on Electronics and Telecommunications (PACET) . 1–5
Georgios Tzanos, Christoforos Kachris, and Dimitrios Soudris. 2022 · 2022
Cited alongside, same era.
LightSeq2: Accelerated Training for Transformer-Based Models on GPUs. In SC22: International Conference for High Performance Computing, Networking, Storage and Analysis . 1–14
Xiaohui Wang, Yang Wei, Ying Xiong, Guyue Huang, Xian Qian, Yufei Ding, Mingxuan Wang, and Lei Li. 2022 · 2022
Cited alongside, same era.
Energon: Toward Efficient Acceleration of Transformers Using Dynamic Sparse Attention
Zhe Zhou, Junlin Liu, Zhenyu Gu, and Guangyu Sun. 2023 · 2022
Cited alongside, same era.
Transformer-OPU: An FPGA-based Overlay Processor for Transformer Networks. In 2023 IEEE 31st Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) . 221–221
Yueyin Bai, Hao Zhou, Keqing Zhao, Jianli Chen, Jun Yu, and Kun Wang. 2023 · 2023
Cited alongside, same era.
Exponentially Faster Language Modelling
Peter Belcak and Roger Wattenhofer. 2023 · 2023
Cited alongside, same era.
Shrihari Sridharan, Jacob R. Stevens, Kaushik Roy, and Anand Raghunathan. 2023 · 2023
Later among the works it cites.
Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation
Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui. 2023 · 2023
Later among the works it cites.
Inference with Reference: Lossless Acceleration of Large Language Models
Nan Yang, Tao Ge, Liang Wang, Binxing Jiao, Daxin Jiang, Linjun Yang, Rangan Majumder, and Furu Wei. 2023 · 2023
Later among the works it cites.
Transformer-based models and hardware acceleration analysis in autonomous driving: A survey
Juan Zhong, Zheng Liu, and Xi Chen. 2023 · 2023
Later among the works it cites.
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, Minsik Cho, Carlo C Del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar. 2024 · 2024
Closest in time.
A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODE
Ikumi Okubo, Keisuke Sugiura, and Hiroki Matsutani. 2024 · 2024
Closest in time.