Fetching the paper…
Reading the bibliography…
In this paper, we propose LoopLynx, a scalable dataflow architecture for efficient LLM inference that optimizes FPGA usage through a hybrid spatial-temporal design.
Attention is all you need
Vaswani et al · 2017
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi et al · 2019
Earlier work this paper cites.
Ftrans: energy-efficient acceleration of transformers using fpga
Bingbing Li et al · 2020
Earlier work this paper cites.
Npe: An fpga-based overlay processor for natural language processing
Hamza Khan et al · 2021
Earlier work this paper cites.
Hardware acceleration of fully quantized BERT for efficient natural language processing
Zejian Liu et al · 2021
Earlier work this paper cites.
Accommodating transformer onto fpga: Coupling the balanced model compression and fpga-implementation optimization
Panjie Qi et al · 2021
Earlier work this paper cites.
Accelerating framework of transformer by hardware design and model compression co-optimization
Panjie Qi et al · 2021
Earlier work this paper cites.
Accelerating transformer-based deep learning models on fpgas using column balanced block pruning
Hongwu Peng et al · 2021
Cited alongside, same era.
Efficient methods for mapping neural machine translator on fpgas
Qin Li et al · 2021
Cited alongside, same era.
DFX: A low-latency multi-fpga appliance for accelerating transformer-based text generation
Seongmin Hong et al · 2022
Cited alongside, same era.
Trac: Compilation-based design of transformer accelerators for fpgas
Patrick Plagwitz et al · 2022
Cited alongside, same era.
Llm-empowered chatbots for psychiatrist and patient simulation: Application and evaluation
Siyuan Chen et al · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
Reiner Pope et al · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron et al · 2023
Later among the works it cites.
SmoothQuant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao et al · 2023
Later among the works it cites.
Flightllm: Efficient large language model inference with a complete mapping flow on fpgas
Shulin Zeng et al · 2024
Later among the works it cites.
Understanding the potential of fpga-based spatial acceleration for large language model inference
Hongzheng Chen et al · 2024
Later among the works it cites.
https://www.amd.com/en/products/accelerators/alveo/u50/a-u50-p00g-pq-g.html
AMD Alveo U50 Card · 2024
Later among the works it cites.
https://developer.nvidia.com/system-management-interface
NVIDIA System Management Interface · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Later among the works it cites.