Fetching the paper…
Reading the bibliography…
Transformers have transformed the field of natural language processing.
A. Vaswani, N. Shazeer, N. Parmar et al. , “Attention is all you need,” in Proc. of NeurIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
“NVIDIA Tesla V100 GPU Architecture,” NVIDIA, Tech. Rep., 2017
2017
Earlier work this paper cites.
N. P. Jouppi, C. Young, N. Patil et al. , “In-datacenter performance analysis of a tensor processing unit,” in Proc. of ISCA , 2017, pp. 1–12
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. of NAACL-HLT , 2019, pp. 4171–4189
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Q. Guo, X. Qiu, P. Liu et al. , “Star-Transformer,” in NAACL-HLT , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
G. Du, C. Tian, Z. Li et al. , “Efficient softmax hardware architecture for deep neural networks,” in Proc. of GLSVLSI , 2019, pp. 75–80
2019
Cited alongside, same era.
R. Venkatesan et al. , “MAGNet: A modular aaccelerator generator for neural networks.” in Proc. of ICCAD , 2019, pp. 1–8
2019
Cited alongside, same era.
T. Wolf et al. , “Huggingface’s transformers: State-of-the-art natural language processing,” ArXiv , pp. arXiv–1910, 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
G. Prato, E. Charlaix, and M. Rezagholizadeh, “Fully quantized transformer for machine translation,” in Proc. of EMNLP , 2020
2020
Later among the works it cites.
Y. Lin, Y. Li, T. Liu et al. , “Towards fully 8-bit integer inference for the transformer model,” in Proc. of IJCAI , 2020, pp. 3759–3765
2020
Later among the works it cites.
Y. Gao, W. Liu, and F. Lombardi, “Design and implementation of an approximate softmax layer for deep neural networks,” in Proc. of ISCAS . IEEE, 2020, pp. 1–5
2020
Later among the works it cites.
T. J. Ham, S. J. Jung, S. Kim et al. , “Aˆ3: Accelerating attention mechanisms in neural networks with approximation,” in Proc. of HPCA . IEEE, 2020, pp. 328–341
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2020
Cited alongside, same era.
[Online]. Available: https://www.synopsys.com/designware-ip.html
Cited in the paper.
2020
Later among the works it cites.
D. Zhu, S. Lu, M. Wang, J. Lin, and Z. Wang, “Efficient precision-adjustable architecture for softmax function in deep learning,” IEEE Transactions on Circuits and Systems II: Express Briefs , 2020
2020
Later among the works it cites.