Fetching the paper…
Reading the bibliography…
Transformers are widely used in state-of-the-art machine translation, but the key to their success is still unknown.
Joseph Enguehard, Dan Busbridge, Vitalii Zhelezniak, and Nils Hammerla. 2019 · 1910
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
Learning grammatical structure with echo state networks
Matthew H. Tong, Adam D. Bickett, Eric M. Christiansen, and Garrison W. Cottrell. 2007 · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Y. Bengio. 2010 · 2010
Earlier work this paper cites.
An echo state network with working memories for probabilistic language modeling
Yukinori Homma and Masafumi Hagiwara. 2013 · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
On the saddle point problem for non-convex optimization
Razvan Pascanu, Yann N. Dauphin, Surya Ganguli, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Transfer learning for low-resource neural machine translation
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Pre-wiring and pre-training: What does a neural network need to learn truly general identity rules?
Raquel Alhama and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Findings of the 2018 conference on machine translation (WMT18)
Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Philipp Koehn, and Christof Monz. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Marian: Fast neural machine translation in C++
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, André F. T. Martins, and Alexandra Birch. 2018 · 2018
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
No training required: Exploring random encoders for sentence classification
John Wieting and Douwe Kiela. 2019 · 2019
Later among the works it cites.
In neural machine translation, what does transfer learning transfer?
Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, and Rico Sennrich. 2020 · 2020
Closest in time.
Edinburgh’s submissions to the 2020 machine translation efficiency task
Nikolay Bogoychev, Roman Grundkiewicz, Alham Fikri Aji, Maximiliana Behnke, Kenneth Heafield, Sidharth Kashyap, Emmanouil-Ioannis Farsarakis, and Mateusz Chudyk. 2020 · 2020
Closest in time.
Echo state neural machine translation
Ankush Garg, Yuan Cao, and Qi Ge. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Freezing subnetworks to analyze domain adaptation in neural machine translation
Brian Thompson, Huda Khayrallah, Antonios Anastasopoulos, Arya D. McCarthy, Kevin Duh, Rebecca Marvin, Paul McNamee, Jeremy Gwinnup, Tim Anderson, and Philipp Koehn. 2018 · 2018
Cited alongside, same era.
Can you tell me how to get past sesame street? sentence-level pretraining beyond language modeling
Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
Language modeling teaches you more than translation does: Lessons learned through auxiliary syntactic task analysis
Kelly Zhang and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
On the use of BERT for neural machine translation
Stephane Clinchant, Kweon Woo Jung, and Vassilina Nikoulina. 2019 · 2019
Cited alongside, same era.
From research to production and back: Ludicrously fast neural machine translation
Young Jin Kim, Marcin Junczys-Dowmunt, Hany Hassan, Alham Fikri Aji, Kenneth Heafield, Roman Grundkiewicz, and Nikolay Bogoychev. 2019 · 2019
Cited alongside, same era.
Echo state networks for named entity recognition
Rajkumar Ramamurthy, Robin Stenzel, Rafet Sifa, Anna Ladi, and Christian Bauckhage. 2019 · 2019
Cited alongside, same era.
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2020 · 2020
Closest in time.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler. 2020 · 2020
Closest in time.
Fixed encoder self-attention patterns in transformer-based machine translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann. 2020 · 2020
Closest in time.
Karthik A. Sankararaman, Soham De, Zheng Xu, W. Ronny Huang, and Tom Goldstein. 2020 · 2020
Closest in time.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2020 · 2020
Closest in time.
Hanxiao Liu, Zihang Dai, David R. So, and Quoc V. Le. 2021 · 2021
Closest in time.