Fetching the paper…
Reading the bibliography…
The self-attention mechanism distinguishes transformer-based large language models (LLMs) apart from convolutional and recurrent neural networks.
An exploration of softmax alternatives belonging to the spherical loss family
A. Brébisson and P. Vincent · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity et al · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin et al · 2018
Earlier work this paper cites.
Enhanced transformer model for data-to-text generation
Gong et al · 2019
Earlier work this paper cites.
Efficient softmax hardware architecture for deep neural networks
G. Du et al · 2019
Earlier work this paper cites.
Toward an open-source digital flow: First learnings from the openroad project
Tutu Ajayi, Vidya A Chhabria, Mateus Fogaça, Soheil Hashemi, Abdelrahman Hosny, Andrew B Kahng, Minsoo Kim, Jeongsup Lee, Uday Mallappa, Marina Neseem, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown et al · 2020
Earlier work this paper cites.
A3: Accelerating attention mechanisms in neural networks with approximation
T. Ham et al · 2020
Earlier work this paper cites.
Exploring alternatives to softmax function
E. Banerjee et al · 2020
Earlier work this paper cites.
Vivit: A video vision transformer
A. Arnab et al · 2021
Earlier work this paper cites.
Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks
T. Ham et al · 2021
Earlier work this paper cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
H. Wang et al · 2021
Cited alongside, same era.
Softermax: Hardware/software co-design of an efficient softmax for transformers
J. Stevens et al · 2021
Cited alongside, same era.
Post-training quantization for vision transformer
Z. Liu et al · 2021
Cited alongside, same era.
Mixed precision quantization of transformer language models for speech recognition
J. Xu et al · 2021
Cited alongside, same era.
Mind mappings: enabling efficient algorithm-accelerator mapping space search
Kartik Hegde, Po-An Tsai, Sitao Huang, Vikas Chandra, Angshuman Parashar, and Christopher W Fletcher · 2021
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Base-2 softmax function: Suitability for training and efficient hardware implementation
Y. Zhang et al · 2022
Later among the works it cites.
Nn-lut: neural approximation of non-linear operations for efficient transformer inference
Y. Joonsang et al · 2022
Later among the works it cites.
A. Jiang et al · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
H. Touvron et al · 2023
Later among the works it cites.
Quip: 2-bit quantization of large language models with guarantees
J. Chee et al · 2023
Later among the works it cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Frantar et al · 2022
Cited alongside, same era.
Sparse attention acceleration with synergistic in-memory pruning and on-chip recomputation
A. Yazdanbakhsh et al · 2022
Cited alongside, same era.
Dota: detect and omit weak attentions for scalable transformer acceleration
Z. Qu et al · 2022
Cited alongside, same era.
Transpim: A memory-based acceleration via software-hardware co-design for transformer
M. Zhou et al · 2022
Cited alongside, same era.
Dfx: A low-latency multi-fpga appliance for accelerating transformer-based text generation
S. Hong et al · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao et al · 2022
Cited alongside, same era.
F. Frantar et al · 2023
Later among the works it cites.
16.2 a 28nm 53.8tops/w 8b sparse transformer accelerator with in-memory butterfly zero skipper for unstructured-pruned nn and cim-based local-attention-reusable engine
S. Liu et al · 2023
Later among the works it cites.
Fact: Ffn-attention co-optimized transformer architecture with eager correlation prediction
Y. Qin et al · 2023
Later among the works it cites.
Flashdecoding++: Faster large language model inference on gpus
K. Hong et al · 2023
Later among the works it cites.
Approximate softmax functions for energy-efficient deep neural networks
K. Chen et al · 2023
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
T. Dao · 2023
Later among the works it cites.