Fetching the paper…
Reading the bibliography…
Transformer, BERT and their variants have achieved great success in natural language processing.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Efficient inference with tensorrt
Han Vanholder. 2016 · 2016
Earlier work this paper cites.
A computational model for tensorflow: an introduction
Martín Abadi, Michael Isard, and Derek Gordon Murray. 2017 · 2017
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017b · 2017
Earlier work this paper cites.
TVM: An automated end-to-end optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018 · 2018
Earlier work this paper cites.
Understanding back-translation at scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018 · 2018
Cited alongside, same era.
Diverse beam search for improved description of complex scenes
Ashwin K. Vijayakumar, Michael Cogswell, Ramprasaath R. Selvaraju, Qing Sun, Stefan Lee, David J. Crandall, and Dhruv Batra. 2018 · 2018
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
The curious case of neural text degeneration
Efficient inference for neural machine translation
Yi-Te Hsu, Sarthak Garg, Yi-Hsiu Liao, and Ilya Chatsviorkin. 2020 · 2020
Closest in time.
FastFormers: Highly efficient transformer models for natural language understanding
Young Jin Kim and Hany Hassan. 2020 · 2020
Closest in time.
Pre-training multilingual neural machine translation by leveraging alignment information
Zehui Lin, Xiao Pan, Mingxuan Wang, Xipeng Qiu, Jiangtao Feng, Hao Zhou, and Lei Li. 2020 · 2020
Closest in time.
Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training
Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou. 2020 · 2020
Closest in time.
Towards making the most of BERT in neural machine translation
Jiacheng Yang, Mingxuan Wang, Hao Zhou, Chengqi Zhao, Weinan Zhang, Yong Yu, and Lei Li. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017a
Cited in the paper.
Turbotransformers: an efficient GPU serving system for transformer models
Jiarui Fang, Yang Yu, Chengduo Zhao, and Jie Zhou. 2021 · 2021
Closest in time.