Fetching the paper…
Reading the bibliography…
Non-linear operations such as GELU, Layer normalization, and Softmax are essential yet costly building blocks of Transformer models.
Optimal Curve Fitting With Piecewise Linear Functions
A. Cantoni. 1971 · 1971
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko. 1989 · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
K. Hornik. 1991 · 1991
Earlier work this paper cites.
Neural acceleration for general-purpose approximate programs. In In MICRO
H. Esmaeilzadeh, A. Sampson, L. Ceze, and D. Burger. 2012 · 2012
Earlier work this paper cites.
Neural network-based accelerators for transcendental function approximation. In In GLSVLSI
S. Eldridge, F. Raudies, D. Zou, and A. Joshi. 2014 · 2014
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text. In In EMNLP
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang. 2016 · 2016
Earlier work this paper cites.
A high-performance deeply pipelined architecture for elementary transcendental function evaluation. In In ICCD
J. Chen and X. Liu. 2017 · 2017
Earlier work this paper cites.
The expressive power of neural networks: a view from the width. In In NeurIPS
Z. Lu et al · 2017
Earlier work this paper cites.
Attention is all you need. In In NeurIPS
A. Vaswani et al · 2017
Cited alongside, same era.
AXNet: ApproXimate computing using an end-to-end trainable neural network. In In ICCAD
Z. Peng et al · 2018
Cited alongside, same era.
Efficient 8-bit quantization of transformer neural machine language translation model. In In ICML workshop
Bhandare et al · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In In NAACL-HLT
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. 2019 · 2019
Cited alongside, same era.
RoBERTa: A robustly optimized bert pretraining approach
Y. Liu et al · 2019
Cited alongside, same era.
GPT-3: Its nature, scope, limits, and consequences
L. Floridi and M. Chiriatti. 2020 · 2020
Later among the works it cites.
MobileBERT: a compact task-agnostic bert for resource-limited devices. In In ACL
Z. Sun et al · 2020
Later among the works it cites.
http://nvdla.org/primer.html
NVIDIA Deep Learning Accelerator · 2021
Closest in time.
Sparsity-Aware and Re-configurable NPU Architecture for Samsung Flagship Mobile SoC. In In ISCA
J.-W. Jang et al · 2021
Closest in time.
I-BERT: Integer-only BERT Quantizatio. In In ICML
S. Kim et al · 2021
Closest in time.
9.5 A 6K-MAC Feature-Map-Sparsity-Aware Neural Processing Unit in 5nm Flagship Mobile SoC. In In ISSCC
J.-S. Park et al · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Song et al · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding. In In ICLR
A. Wang et al · 2019
Cited alongside, same era.
Q8BERT: Quantized 8bit bert. In In NeurIPS workshop
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat. 2019 · 2019
Cited alongside, same era.
Softermax: Hardware/Software Co-Design of an Efficient Softmax for Transformers. In In DAC
J. R. Stevens et al · 2021
Closest in time.
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning. In In HPCA
H. Wang, Z. Zhang, and S. Han. 2021 · 2021
Closest in time.