Fetching the paper…
Reading the bibliography…
Transformer-based Large Language Models (LLMs) have made a significant impact on various domains.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Large language models in medicine
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. 2023 · 1940
Earlier work this paper cites.
A 3 : Accelerating Attention Mechanisms in Neural Networks with Approximation
Tae Jun Ham, Sung Jun Jung, Seonghak Kim, Young H. Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W. Lee, and Deog-Kyoon Jeong. 2020 · 2002
Earlier work this paper cites.
FTRANS: Energy-Efficient Acceleration of Transformers using FPGA
Bingbing Li, Santosh Pandey, Haowen Fang, Yanjun Lyv, Ji Li, Jieyang Chen, Mimi Xie, Lipeng Wan, Hang Liu, and Caiwen Ding. 2020 · 2007
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
Deep Learning with INT8 Optimization on Xilinx Devices
2017 · 2017
Earlier work this paper cites.
Angel-eye: A complete design flow for mapping CNN onto embedded FPGA
Kaiyuan Guo, Lingzhi Sui, Jiantao Qiu, Jincheng Yu, Junbin Wang, Song Yao, Song Han, Yu Wang, and Huazhong Yang. 2017 · 2017
Earlier work this paper cites.
Bandwidth and locality aware task-stealing for manycore architectures with bandwidth-asymmetric memory
Han Zhao, Quan Chen, Yuxian Qiu, Ming Wu, Yao Shen, Jingwen Leng, Chao Li, and Minyi Guo. 2018 · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2020
Earlier work this paper cites.
Model compression and hardware acceleration for neural networks: A comprehensive survey
Lei Deng, Guoqi Li, Song Han, Luping Shi, and Yuan Xie. 2020 · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Earlier work this paper cites.
Matraptor: A sparse-sparse matrix multiplication accelerator based on row-wise product. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 766–780
Nitish Srivastava, Hanchen Jin, Jie Liu, David Albonesi, and Zhiru Zhang. 2020 · 2020
Earlier work this paper cites.
Shuhai: Benchmarking high bandwidth memory on fpgas. In 2020 IEEE 28th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) . IEEE, 111–119
Zeke Wang, Hongjing Huang, Jie Zhang, and Gustavo Alonso. 2020 · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Earlier work this paper cites.
Xilinx Board Utility Tool
2022 · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Nvidia a100 tensor core gpu: Performance and innovation
Jack Choquette, Wishwesh Gandhi, Olivier Giroux, Nick Stam, and Ronny Krashinsky. 2021 · 2021
Earlier work this paper cites.
ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural Networks. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) . 692–705
Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W. Lee. 2021 · 2021
Cited alongside, same era.
Sanger: A Co-Design Framework for Enabling Sparse Attention Using Reconfigurable Architecture. In MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture (Virtual Event, Greece) (MICRO ’21) . Association for Computing Machinery, New York, NY, USA, 977–991
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang. 2021 · 2021
Cited alongside, same era.
Spatten: Efficient sparse attention architecture with cascade token and head pruning. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 97–110
Hanrui Wang, Zhekai Zhang, and Song Han. 2021 · 2021
Cited alongside, same era.
Alveo U280 Data Center Accelerator Card Data Sheet
Xilinx. 2021 · 2021
Cited alongside, same era.
LLM-empowered Chatbots for Psychiatrist and Patient Simulation: Application and Evaluation
Siyuan Chen, Mengyue Wu, Kenny Q Zhu, Kunyao Lan, Zhiling Zhang, and Lyuchun Cui. 2023b · 2023
Later among the works it cites.
RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset
Together Computer. 2023 · 2023
Later among the works it cites.
ChatLaw: Open-Source Legal Large Language Model with Integrated External Knowledge Bases
Jiaxi Cui, Zongjian Li, Yang Yan, Bohua Chen, and Li Yuan. 2023 · 2023
Later among the works it cites.
Diffuser: efficient transformers with multi-hop attention diffusion for long sequences. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 12772–12780
Aosong Feng, Irene Li, Yuang Jiang, and Rex Ying. 2023 · 2023
Later among the works it cites.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
FracBNN: Accurate and FPGA-efficient binary neural networks with fractional activations. In The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . 171–182
Yichi Zhang, Junhao Pan, Xinheng Liu, Hongzheng Chen, Deming Chen, and Zhiru Zhang. 2021 · 2021
Cited alongside, same era.
Learning n: m fine-grained structured sparse neural networks from scratch
Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li. 2021 · 2021
Cited alongside, same era.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022a · 2022
Cited alongside, same era.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022b · 2022
Cited alongside, same era.
Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design. In 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) . 599–615
Hongxiang Fan, Thomas Chau, Stylianos I. Venieris, Royson Lee, Alexandros Kouris, Wayne Luk, Nicholas D. Lane, and Mohamed S. Abdelfattah. 2022 · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022 · 2022
Cited alongside, same era.
N3H-core: Neuron-designed neural network accelerator via FPGA-based heterogeneous computing cores. In Proceedings of the 2022 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . 112–122
Yu Gong, Zhihan Xu, Zhezhi He, Weifeng Zhang, Xiaobing Tu, Xiaoyao Liang, and Li Jiang. 2022 · 2022
Cited alongside, same era.
DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation. In 2022 IEEE Hot Chips 34 Symposium (HCS) . 1–17
Seongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee, Minsub Kim, Dongsoo Lee, and Joo-Young Kim. 2022 · 2022
Cited alongside, same era.
Elias Frantar and Dan Alistarh. 2023 · 2023
Later among the works it cites.
Cheonsu Jeong. 2023 · 2023
Later among the works it cites.
FLAT: An Optimized Dataflow for Mitigating Attention Bottlenecks. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (Vancouver, BC, Canada) (ASPLOS 2023) . Association for Computing Machinery, New York, NY, USA, 295–310
Sheng-Chun Kao, Suvinay Subramanian, Gaurav Agrawal, Amir Yazdanbakhsh, and Tushar Krishna. 2023 · 2023
Later among the works it cites.
SqueezeLLM: Dense-and-Sparse Quantization
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer. 2023 · 2023
Later among the works it cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles . 611–626
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Later among the works it cites.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023 · 2023
Later among the works it cites.
A Comprehensive Overview of Large Language Models
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Nick Barnes, and Ajmal S. Mian. 2023 · 2023
Later among the works it cites.
FACT: FFN-Attention Co-Optimized Transformer Architecture with Eager Correlation Prediction. In Proceedings of the 50th Annual International Symposium on Computer Architecture (Orlando, FL, USA) (ISCA ’23) . Association for Computing Machinery, New York, NY, USA, Article 22, 14 pages
Yubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao, Xiaolong Yang, Leibo Liu, Shaojun Wei, Yang Hu, and Shouyi Yin. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Practitioners’ Expectations on Code Completion
Chaozheng Wang, Junhao Hu, Cuiyun Gao, Yu Jin, Tao Xie, Hailiang Huang, Zhenyu Lei, and Yuetang Deng. 2023a · 2023
Later among the works it cites.
CTA: Hardware-Software Co-design for Compressed Token Attention Mechanism. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . 429–441
Haoran Wang, Haobo Xu, Ying Wang, and Yinhe Han. 2023c · 2023
Later among the works it cites.
COSA: Co-Operative Systolic Arrays for Multi-head Attention Mechanism in Neural Network using Hybrid Data Reuse and Fusion Methodologies. In 2023 60th ACM/IEEE Design Automation Conference (DAC) . IEEE, 1–6
Zhican Wang, Gang Wang, Honglan Jiang, Ningyi Xu, and Guanghui He. 2023b · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning . PMLR, 38087–38099
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023 · 2023
Later among the works it cites.
Versal™ Architecture and Product Data Sheet
Xilinx. 2023 · 2023
Later among the works it cites.