Fetching the paper…
Reading the bibliography…
Asynchronous stochastic gradient descent (SGD) is attractive from a speed perspective because workers do not wait for synchronization.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu. 2011 · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al. 2012 · 2012
Earlier work this paper cites.
Variance reduction for stochastic gradient optimization
Chong Wang, Xi Chen, Alexander J Smola, and Eric P Xing. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Delay-tolerant algorithms for asynchronous distributed online learning
Brendan McMahan and Matthew Streeter. 2014 · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu. 2014 · 2014
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu. 2015 · 2015
Earlier work this paper cites.
Revisiting distributed synchronous sgd
Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz. 2016 · 2016
Earlier work this paper cites.
Model accuracy and runtime tradeoff in distributed deep learning: A systematic study
Suyog Gupta, Wei Zhang, and Fei Wang. 2016 · 2016
Earlier work this paper cites.
Staleness-aware async-sgd for distributed deep learning
Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu. 2016 · 2016
Earlier work this paper cites.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield. 2017 · 2017
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017 · 2017
Cited alongside, same era.
Deep architectures for neural machine translation
Antonio Valerio Miceli Barone, Jindřich Helcl, Rico Sennrich, Barry Haddow, and Alexandra Birch. 2017 · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally. 2017 · 2017
Cited alongside, same era.
Accelerating asynchronous stochastic gradient descent for neural machine translation
Nikolay Bogoychev, Kenneth Heafield, Alham Fikri Aji, and Marcin Junczys-Dowmunt. 2018 · 2018
Later among the works it cites.
Findings of the 2018 conference on machine translation (wmt18)
Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz. 2018 · 2018
Later among the works it cites.
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Niki Parmar, Mike Schuster, Zhifeng Chen, et al. 2018 · 2018
Later among the works it cites.
Slow and stale gradients can win the race: Error-runtime trade-offs in distributed sgd
Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh, Parijat Dube, and Priya Nagpurkar. 2018 · 2018
Later among the works it cites.
Marian: Fast neural machine translation in C++
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2017 · 2017
Cited alongside, same era.
The University of Edinburgh’s neural mt systems for WMT17
Rico Sennrich, Alexandra Birch, Anna Currey, Ulrich Germann, Barry Haddow, Kenneth Heafield, Antonio Valerio Miceli Barone, and Philip Williams. 2017 · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016a
Cited in the paper.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016b
Cited in the paper.
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, André F. T. Martins, and Alexandra Birch. 2018 · 2018
Later among the works it cites.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Later among the works it cites.
Training tips for the transformer model
Martin Popel and Ondřej Bojar. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting bleu scores
Matt Post. 2018 · 2018
Later among the works it cites.
An analysis of the delayed gradients problem in asynchronous sgd
Anand Srinivasan, Ajay Jain, and Parnian Barekatain. 2018 · 2018
Later among the works it cites.