Fetching the paper…
Reading the bibliography…
Stacking non-linear layers allows deep neural networks to model complicated functions, and including residual connections in Transformer layers is beneficial for convergence and performance.
Neutron: An Implementation of the Transformer Translation Model and its Variants
Hongfei Xu and Qiuhui Liu. 2019 · 1903
Earlier work this paper cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George F. Foster, Colin Cherry, Wolfgang Macherey, Zhifeng Chen, and Yonghui Wu. 2019 · 1907
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
GLU variants improve transformer
Noam Shazeer. 2020 · 2002
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Parallel data, tools and interfaces in opus
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber. 2015 · 2015
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir. 2016 · 2016
Earlier work this paper cites.
Multi-way, multilingual neural machine translation with a shared attention mechanism
Orhan Firat, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
benefits of depth in neural networks
Matus Telgarsky. 2016 · 2016
Earlier work this paper cites.
Deep recurrent models with fast-forward connections for neural machine translation
Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, and Wei Xu. 2016 · 2016
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2017 · 2017
Earlier work this paper cites.
When and why are deep networks better than shallow ones?
Hrushikesh Mhaskar, Qianli Liao, and Tomaso Poggio. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Deep neural machine translation with linear associative unit
Mingxuan Wang, Zhengdong Lu, Jie Zhou, and Qun Liu. 2017 · 2017
Cited alongside, same era.
Training deeper neural machine translation models with transparent attention
Ankur Bapna, Mia Chen, Orhan Firat, Yuan Cao, and Yonghui Wu. 2018 · 2018
Cited alongside, same era.
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, and Macduff Hughes. 2018 · 2018
Cited alongside, same era.
Improving deep transformer with depth-scaled initialization and merged attention
Biao Zhang, Ivan Titov, and Rico Sennrich. 2019 · 2019
Later among the works it cites.
Highway transformer: Self-gating enhanced self-attentive networks
Yekun Chai, Shuo Jin, and Xinwen Hou. 2020 · 2020
Closest in time.
Improving transformer optimization through better initialization
Xiao Shi Huang, Felipe Perez, Jimmy Ba, and Maksims Volkovs. 2020 · 2020
Closest in time.
Shallow-to-deep training for neural machine translation
Bei Li, Ziyang Wang, Hui Liu, Yufan Jiang, Quan Du, Tong Xiao, Huizhen Wang, and Jingbo Zhu. 2020 · 2020
Closest in time.
Multiscale collaborative deep models for neural machine translation
Xiangpeng Wei, Heng Yu, Yue Hu, Yue Zhang, Rongxiang Weng, and Weihua Luo. 2020 · 2020
Closest in time.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploiting deep representations for neural machine translation
Zi-Yi Dou, Zhaopeng Tu, Xing Wang, Shuming Shi, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
Dense information flow for neural machine translation
Yanyao Shen, Xu Tan, Di He, Tao Qin, and Tie-Yan Liu. 2018 · 2018
Cited alongside, same era.
Multi-layer representation fusion for neural machine translation
Qiang Wang, Fuxue Li, Tong Xiao, Yanyang Li, Yinqiao Li, and Jingbo Zhu. 2018 · 2018
Cited alongside, same era.
Deep layer aggregation
Fisher Yu, Dequan Wang, Evan Shelhamer, and Trevor Darrell. 2018 · 2018
Cited alongside, same era.
Closest in time.
Improving massively multilingual neural machine translation and zero-shot translation
Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. 2020 · 2020
Closest in time.
Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation
Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah Smith. 2021 · 2021
Closest in time.
Learning light-weight translation models from deep transformer
Bei Li, Ziyang Wang, Hui Liu, Quan Du, Tong Xiao, Chunliang Zhang, and Jingbo Zhu. 2021 · 2021
Closest in time.
Luna: Linear unified nested attention
Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma, and Luke Zettlemoyer. 2021 · 2021
Closest in time.
Delight: Deep and light-weight transformer
Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2021 · 2021
Closest in time.
Probing word translations in the transformer and trading decoder for encoder layers
Hongfei Xu, Josef van Genabith, Qiuhui Liu, and Deyi Xiong. 2021c · 2021
Closest in time.
Optimizing deep transformers for chinese-thai low-resource translation
Wenjie Hao, Hongfei Xu, Lingling Mu, and Hongying Zan. 2022 · 2022
Closest in time.
What works and doesn’t work, a deep decoder for neural machine translation
Zuchao Li, Yiran Wang, Masao Utiyama, Eiichiro Sumita, Hai Zhao, and Taro Watanabe. 2022b · 2022
Closest in time.
Optimizing deeper transformers on small datasets
Peng Xu, Dhruv Kumar, Wei Yang, Wenjie Zi, Keyi Tang, Chenyang Huang, Jackie Chi Kit Cheung, Simon J.D. Prince, and Yanshuai Cao. 2021d · 2089
Closest in time.