Fetching the paper…
Reading the bibliography…
PipeDream is a Deep Neural Network(DNN) training system for GPUs that parallelizes computation by pipelining execution across multiple machines.
On bayesian methods for seeking the extremum
J. Močkus · 1975
Earlier work this paper cites.
A bridging model for parallel computation
L. G. Valiant · 1990
Earlier work this paper cites.
Optimization of collective communication operations in mpich
R. Thakur, R. Rabenseifner, and W. Gropp · 2005
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Pipelined Back-propagation for Context-dependent Deep Neural Networks
X. Chen, A. Eversole, G. Li, D. Yu, and F. Seide · 2012
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Distributed graphlab: A framework for machine learning in the cloud
Y. Low, J. Gonzalez, A. Kyrola, D. Bickson, C. Guestrin, and J. M. Hellerstein · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P. B. Gibbons, G. A. Gibson, G. Ganger, and E. P. Xing · 2013
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
T. M. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman · 2014
Earlier work this paper cites.
Exploiting bounded staleness to speed up big data analytics
H. Cui, J. Cipar, Q. Ho, J. K. Kim, S. Lee, A. Kumar, J. Wei, W. Dai, G. R. Ganger, P. B. Gibbons, et al · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
A convolutional neural network for modelling sentences
N. Kalchbrenner, E. Grefenstette, and P. Blunsom · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Automating model search for large scale machine learning
E. R. Sparks, A. Talwalkar, D. Haas, M. J. Franklin, M. I. Jordan, and T. Kraska · 2015
Later among the works it cites.
Sequence to sequence-video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Later among the works it cites.
Lightlda: Big topic models on modest computer clusters
J. Yuan, F. Gao, Q. Ho, W. Dai, J. Wei, X. Zheng, E. P. Xing, T. Liu, and W. Ma · 2015
Later among the works it cites.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Later among the works it cites.
Revisiting distributed synchronous sgd
J. Chen, R. Monga, S. Bengio, and R. Jozefowicz · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
One weird trick for parallelizing convolutional neural networks
A. Krizhevsky · 2014
Cited alongside, same era.
On model parallelization and scheduling strategies for distributed machine learning
S. Lee, J. K. Kim, X. Zheng, Q. Ho, G. A. Gibson, and E. P. Xing · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Cited alongside, same era.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Geeps: Scalable deep learning on distributed gpus with a gpu-specialized parameter server
H. Cui, H. Zhang, G. R. Ganger, P. B. Gibbons, and E. P. Xing · 2016
Later among the works it cites.
Omnivore: An optimizer for multi-device deep learning on cpus and gpus
S. Hadjis, C. Zhang, I. Mitliagkas, D. Iter, and C. Ré · 2016
Later among the works it cites.
STRADS: a distributed framework for scheduled model parallel machine learning
J. K. Kim, Q. Ho, S. Lee, X. Zheng, W. Dai, G. A. Gibson, and E. P. Xing · 2016
Later among the works it cites.
CNTK: Microsoft’s open-source deep-learning toolkit
F. Seide and A. Agarwal · 2016
Later among the works it cites.
Bringing HPC Techniques to Deep Learning, 2017
Baidu Inc · 2017
Later among the works it cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Later among the works it cites.
Device placement optimization with reinforcement learning
A. Mirhoseini, H. Pham, Q. Le, M. Norouzi, S. Bengio, B. Steiner, Y. Zhou, N. Kumar, R. Larsen, and J. Dean · 2017
Later among the works it cites.
Meet Horovod: Uber’s Open Source Distributed Deep Learning Framework for TensorFlow, 2017
Uber Technologies Inc · 2017
Later among the works it cites.
Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters
H. Zhang, Z. Zheng, S. Xu, W. Dai, Q. Ho, X. Liang, Z. Hu, J. Wei, P. Xie, and E. P. Xing · 2017
Later among the works it cites.