Fetching the paper…
Reading the bibliography…
This paper presents Rudra, a parameter server based distributed computing framework tuned for training large-scale deep neural networks.
Time, clocks, and the ordering of events in a distributed system
L. Lamport · 1978
Earlier work this paper cites.
Logical time: Capturing causality in distributed systems
M. Raynal and M. Singhal · 1996
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
An architecture for parallel topic models
A. Smola and S. Narayanamurthy · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Y. Bengio · 2012
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng · 2012
Earlier work this paper cites.
Mpi 3.0 standard
M. P. I. Forum · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Adadelta: An adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
Deep learning with cots hpc systems
A. Coates, B. Huval, T. Wang, D. Wu, B. Catanzaro, and N. Andrew · 2013
Cited alongside, same era.
More effective distributed ML via a stale synchronous parallel parameter server
Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P. B. Gibbons, G. A. Gibson, G. Ganger, and E. P. Xing · 2013
Cited alongside, same era.
An empirical study of learning rates in deep neural networks for speech recognition
A. Senior, G. Heigold, M. Ranzato, and K. Yang · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Deep learning with elastic averaging SGD
S. Zhang, A. Choromanska, and Y. LeCun · 2014
Later among the works it cites.
Mariana: Tencent deep learning platform and its applications
Y. Zou, X. Jin, Y. Li, Z. Guo, E. Wang, and B. Xiao · 2014
Later among the works it cites.
The effects of hyperparameters on SGD training of neural networks
T. M. Breuel · 2015
Closest in time.
ImageNet Large Scale Visual Recognition Challenge
O. R. et al · 2015
Closest in time.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
Project Adam: Building an efficient and scalable deep learning training system
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman · 2014
Cited alongside, same era.
Exploiting bounded staleness to speed up big data analytics
H. e. a. Cui · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Closest in time.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Closest in time.
Asynchronous Parallel Stochastic Gradient for Nonconvex Optimization
X. Lian, Y. Huang, Y. Li, and J. Liu · 2015
Closest in time.