Fetching the paper…
Reading the bibliography…
The computational requirements for training deep neural networks (DNNs) have grown to the point that it is now standard practice to parallelize training.
Monte carlo sampling methods using markov chains and their applications
W. K. Hastings · 1970
Earlier work this paper cites.
Worst case analysis of two scheduling algorithms
S. Lam and R. Sethi · 1977
Earlier work this paper cites.
Markov chain Monte Carlo in practice
W. R. Gilks, S. Richardson, and D. Spiegelhalter · 1995
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
https://www.cs.cornell.edu/people/pabo/movie-review-data/ , 2005
Movie review data · 2005
Earlier work this paper cites.
Introduction to Algorithms, Third Edition
T. H. Cormen, C. E. Leiserson, R. L. Rivest, and C. Stein · 2009
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Quincy: Fair scheduling for distributed computing clusters
M. Isard, V. Prabhakaran, J. Currey, U. Wieder, K. Talwar, and A. Goldberg · 2009
Earlier work this paper cites.
Legion: Expressing locality and independence with logical regions
M. Bauer, S. Treichler, E. Slaughter, and A. Aiken · 2012
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, and P. Koehn · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Cited alongside, same era.
Convolutional neural networks for sentence classification
Y. Kim · 2014
Cited alongside, same era.
One weird trick for parallelizing convolutional neural networks
A. Krizhevsky · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. E. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Cited alongside, same era.
Firmament: Fast, centralized cluster scheduling at scale
I. Gog, M. Schwarzkopf, A. Gleave, R. N. M. Watson, and S. Hand · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Densely connected convolutional networks
G. Huang, Z. Liu, and K. Q. Weinberger · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Later among the works it cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Later among the works it cites.
Dependent partitioning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Recurrent neural network regularization
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Cited alongside, same era.
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang · 2015
Cited alongside, same era.
LeNet-5, convolutional neural networks
Y. LeCun · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
https://caffe2.ai , 2016
A New Lightweight, Modular, and Scalable Deep Learning Framework · 2016
Cited alongside, same era.
http://www.statmt.org/wmt16 , 2016
Conference on machine translation · 2016
Cited alongside, same era.
S. Treichler, M. Bauer, R. Sharma, E. Slaughter, and A. Aiken · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, J. Klingner, A. Shah, M. Johnson, X. Liu, L. Kaiser, S. Gouws, Y. Kato, T. Kudo, H. Kazawa, K. Stevens, G. Kurian, N. Patil, W. Wang, C. Young, J. Smith, J. Riesa, A. Rudnick, O. Vinyals, G. Corrado, M. Hughes, and J. Dean · 2016
Later among the works it cites.
https://www.tensorflow.org/performance/benchmarks , 2017
TensorFlow Benchmarks · 2017
Later among the works it cites.
https://pytorch.org , 2017
Tensors and Dynamic neural networks in Python with strong GPU acceleration · 2017
Later among the works it cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Later among the works it cites.
Device placement optimization with reinforcement learning
A. Mirhoseini, H. Pham, Q. V. Le, B. Steiner, R. Larsen, Y. Zhou, N. Kumar, M. Norouzi, S. Bengio, and J. Dean · 2017
Later among the works it cites.
Exploring hidden dimensions in parallelizing convolutional neural networks
Z. Jia, S. Lin, C. R. Qi, and A. Aiken · 2018
Closest in time.
A hierarchical model for device placement
A. Mirhoseini, A. Goldie, H. Pham, B. Steiner, Q. V. Le, and J. Dean · 2018
Closest in time.