Fetching the paper…
Reading the bibliography…
While current state-of-the-art NMT models, such as RNN seq2seq and Transformers, possess a large number of parameters, they are still shallow in comparison to convolutional models used for both text and vision applications.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi. 1994 · 1994
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber, et al. 2001 · 2001
Earlier work this paper cites.
Understanding the exploding gradient problem
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2012 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber. 2015 · 2015
Earlier work this paper cites.
Learning functions: when is deep better than shallow
Hrushikesh Mhaskar, Qianli Liao, and Tomaso Poggio. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Weighted residuals for very deep networks
Falong Shen and Gang Zeng. 2016 · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Matus Telgarsky. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Later among the works it cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma. 2017 · 2017
Later among the works it cites.
Skip connections as effective symmetry-breaking
A. Emin Orhan. 2017 · 2017
Later among the works it cites.
Gradients explode - deep networks are shallow - resnet explained
George Philipp, Dawn Song, and Jaime G. Carbonell. 2017 · 2017
Later among the works it cites.
Svcca: Singular vector canonical correlation analysis for deep understanding and improvement
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep recurrent models with fast-forward connections for neural machine translation
Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, and Wei Xu. 2016 · 2016
Cited alongside, same era.
Deep architectures for neural machine translation
Antonio Valerio Miceli Barone, Jindřich Helcl, Rico Sennrich, Barry Haddow, and Alexandra Birch. 2017 · 2017
Cited alongside, same era.
Massive exploration of neural machine translation architectures
Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc Le. 2017 · 2017
Cited alongside, same era.
Sharp models on dull hardware: Fast and accurate neural machine translation decoding on the cpu
Jacob Devlin. 2017 · 2017
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, and Kilian Q. Weinberger. 2016a
Cited in the paper.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger. 2016b
Cited in the paper.
Later among the works it cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Deep neural machine translation with linear associative unit
Mingxuan Wang, Zhengdong Lu, Jie Zhou, and Qun Liu. 2017 · 2017
Later among the works it cites.
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Niki Parmar, Mike Schuster, Zhifeng Chen, Yonghui Wu, and Macduff Hughes. 2018 · 2018
Closest in time.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Closest in time.