Fetching the paper…
Reading the bibliography…
The Transformer architecture has led to significant gains in machine translation.
Learning Representations by Back-propagating Errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. 1986 · 1986
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Statistical phrase-based translation
Philipp Koehn, Franz J. Och, and Daniel Marcu. 2003 · 2003
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
Cache-based document-level statistical machine translation
Zhengxian Gong, Min Zhang, and Guodong Zhou. 2011 · 2011
Earlier work this paper cites.
Docent: A document-level decoder for phrase-based statistical machine translation
Christian Hardmeier, Sara Stymne, Jörg Tiedemann, and Joakim Nivre. 2013 · 2013
Earlier work this paper cites.
Recurrent continuous translation models
Nal Kalchbrenner and Phil Blunsom. 2013 · 2013
Earlier work this paper cites.
Overcoming the curse of sentence length for neural machine translation using automatic segmentation
Jean Pouget-Abadie, Dzmitry Bahdanau, Bart van Merriënboer, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Document-level machine translation with word vector models
Eva Martínez Garcia, Cristina España-Bonet, and Lluís Màrquez. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Memory-augmented neural machine translation
Yang Feng, Shiyue Zhang, Andi Zhang, Dong Wang, and Andrew Abel. 2017 · 2017
Earlier work this paper cites.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Neural machine translation with extended context
Jörg Tiedemann and Yves Scherrer. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Contextual handling in neural machine translation: Look behind, ahead and on both sides
Ruchit Agrawal, Marco Turchi, and Matteo Negri. 2018 · 2018
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross. 2018 · 2018
Cited alongside, same era.
Has machine translation achieved human parity? a case for document-level evaluation
Samuel Läubli, Rico Sennrich, and Martin Volk. 2018 · 2018
Context-aware monolingual repair for neural machine translation
Elena Voita, Rico Sennrich, and Ivan Titov. 2019a · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
ETC: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2020
Later among the works it cites.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
A simple and effective unified encoder for document-level machine translation
Shuming Ma, Dongdong Zhang, and Ming Zhou. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Document context neural machine translation with memory networks
Sameen Maruf and Gholamreza Haffari. 2018 · 2018
Cited alongside, same era.
Document-level neural machine translation with hierarchical attention networks
Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018 · 2018
Cited alongside, same era.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018 · 2018
Cited alongside, same era.
Improving the transformer translation model with document-level context
Jiacheng Zhang, Huanbo Luan, Maosong Sun, Feifei Zhai, Jingfang Xu, Min Zhang, and Yang Liu. 2018 · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Wen-tau Yih, Sinong Wang, and Jie Tang. 2020 · 2020
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. 2020 · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020 · 2020
Later among the works it cites.
Long-short term masking transformer: A simple but effective baseline for document-level neural machine translation
Pei Zhang, Boxing Chen, Niyu Ge, and Kai Fan. 2020 · 2020
Later among the works it cites.
Towards making the most of context in neural machine translation
Zaixiang Zheng, Xiang Yue, Shujian Huang, Jiajun Chen, and Alexandra Birch. 2020 · 2020
Later among the works it cites.
G-transformer for document-level machine translation
Guangsheng Bao, Yue Zhang, Zhiyang Teng, Boxing Chen, and Weihua Luo. 2021 · 2021
Later among the works it cites.
Diverse pretrained context encodings improve document translation
Domenic Donato, Lei Yu, and Chris Dyer. 2021 · 2021
Later among the works it cites.
Fast and accurate neural machine translation with translation memory
Qiuxiang He, Guoping Huang, Qu Cui, Li Li, and Lemao Liu. 2021 · 2021
Later among the works it cites.
Document-level neural machine translation with associated memory network
Shu Jiang, Rui Wang, Zuchao Li, Masao Utiyama, Kehai Chen, Eiichiro Sumita, Hai Zhao, and Bao liang Lu. 2021 · 2021
Later among the works it cites.
∞ \infty -former: Infinite memory transformer
Pedro Henrique Martins, Zita Marinho, and André F. T. Martins. 2021 · 2021
Later among the works it cites.