Fetching the paper…
Reading the bibliography…
This article describes our experiments in neural machine translation using the recent Tensor2Tensor framework and the Transformer sequence-to-sequence model (Vaswani et al., 2017).
BLEU: a Method for Automatic Evaluation of Machine Translation
Papineni, Kishore, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
The Joy of Parallelism with CzEng 1.0
Bojar, Ondřej, Zdeněk Žabokrtský, Ondřej Dušek, Petra Galuščáková, Martin Majliš, David Mareček, Jiří Maršík, Michal Novák, Martin Popel, and Aleš Tamchyna · 2012
Earlier work this paper cites.
Stochastic Gradient Descent Tricks , pages 421–436
Bottou, Léon · 2012
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, Alex · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, Dzmitry, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe, Sergey and Christian Szegedy · 2015
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Popović, Maja · 2015
Earlier work this paper cites.
CzEng 1.6: Enlarged Czech-English Parallel Corpus with Processing Tools Dockered
Bojar, Ondřej, Ondřej Dušek, Tom Kocmi, Jindřich Libovický, Michal Novák, Martin Popel, Roman Sudarikov, and Dušan Variš · 2016
Earlier work this paper cites.
Optimization Methods for Large-Scale Machine Learning
Bottou, L., F. E. Curtis, and J. Nocedal · 2016
Cited alongside, same era.
Fully Character-Level Neural Machine Translation without Explicit Segmentation
Lee, Jason, Kyunghyun Cho, and Thomas Hofmann · 2016
Cited alongside, same era.
Layer Normalization
Lei Ba, J., J. R. Kiros, and G. E. Hinton · 2016
Cited alongside, same era.
Neural Machine Translation of Rare Words with Subword Units
Sennrich, Rico, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Wu, Yonghui, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, Elad, Itay Hubara, and Daniel Soudry · 2017
Later among the works it cites.
Three Factors Influencing Minima in SGD
Jastrzebski, Stanislaw, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos J. Storkey · 2017
Later among the works it cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Keskar, Nitish Shirish, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Later among the works it cites.
A Bayesian Perspective on Generalization and Stochastic Gradient Descent
Smith, Samuel L. and Quoc V. Le · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Findings of the 2017 Conference on Machine Translation (WMT17)
Bojar, Ondřej, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi · 2017
Cited alongside, same era.
Overview of the IWSLT 2017 Evaluation Campaign
Cettolo, Mauro, Marcello Federico, Luisa Bentivogli, Jan Niehues, Sebastian Stüker, Katsuhito Sudoh, Koichiro Yoshino, and Christian Federmann · 2017
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Goyal, Priya, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Results of the WMT17 Metrics Shared Task
Bojar, Ondřej, Yvette Graham, and Amir Kamran
Cited in the paper.
Smith, Samuel L., Pieter-Jan Kindermans, and Quoc V. Le · 2017
Later among the works it cites.
Attention is All you Need
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Scaling SGD Batch Size to 32K for ImageNet Training
You, Yang, Igor Gitman, and Boris Ginsburg · 2017
Later among the works it cites.
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
Shazeer, N. and M. Stern · 2018
Closest in time.