Fetching the paper…
Reading the bibliography…
Deep learning has led to tremendous advancements in the field of Artificial Intelligence.
L. Lamport et al. , “Paxos made simple,” 2001
2001
Earlier work this paper cites.
S. Ghemawat, H. Gobioff, and S.-T. Leung, The Google file system . ACM, 2003, vol. 37, no. 5
2003
Earlier work this paper cites.
R. Thakur, R. Rabenseifner, and W. Gropp, “Optimization of collective communication operations in MPICH,” The International Journal of High Performance Computing Applications , vol. 19, no. 1, pp. 49–66, 2005
2005
Earlier work this paper cites.
T. D. Chandra, R. Griesemer, and J. Redstone, “Paxos made live: an engineering perspective,” in Proceedings of the twenty-sixth annual ACM symposium on Principles of distributed computing . ACM, 2007, pp. 398–407
2007
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” Citeseer, Tech. Rep., 2009
2009
Earlier work this paper cites.
K. Shvachko, H. Kuang, S. Radia, and R. Chansler, “The hadoop distributed file system,” in Mass storage systems and technologies (MSST), 2010 IEEE 26th symposium on . Ieee, 2010, pp. 1–10
2010
Earlier work this paper cites.
L. Bottou, “Large-scale machine learning with stochastic gradient descent,” in Proceedings of COMPSTAT’2010 . Springer, 2010, pp. 177–186
2010
Earlier work this paper cites.
P. Hunt, M. Konar, F. P. Junqueira, and B. Reed, “ZooKeeper: Wait-free Coordination for Internet-scale Systems.” in USENIX annual technical conference , vol. 8, no. 9. Boston, MA, USA, 2010
2010
Earlier work this paper cites.
J. Kreps, N. Narkhede, J. Rao et al. , “Kafka: A distributed messaging system for log processing,” 2011
2011
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in Advances in neural information processing systems , 2011, pp. 693–701
2011
Earlier work this paper cites.
J. Constine, “TechCrunch Article: How Big Is Facebook’s Data? 2.5 Billion Pieces Of Content And 500+ Terabytes Ingested Every Day,” 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
T. Tieleman and G. Hinton, “Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural networks for machine learning , vol. 4, no. 2, pp. 26–31, 2012
2012
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le et al. , “Large scale distributed deep networks,” in Advances in neural information processing systems , 2012, pp. 1223–1231
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
M. Li, D. G. Andersen, and A. Smola, “Distributed delayed proximal gradient methods,” in NIPS Workshop on Optimization for Machine Learning , vol. 3, 2013, p. 3
2013
Earlier work this paper cites.
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” in Advances in neural information processing systems , 2013, pp. 315–323
2013
Earlier work this paper cites.
2014
Cited alongside, same era.
R. Zhang and J. Kwok, “Asynchronous distributed ADMM for consensus optimization,” in International Conference on Machine Learning , 2014, pp. 1701–1709
2014
Cited alongside, same era.
D. Ongaro and J. Ousterhout, “In Search of an Understandable Consensus Algorithm,” in Proceedings of the 2014 USENIX Conference on USENIX Annual Technical Conference , ser. USENIX ATC’14. Berkeley, CA, USA: USENIX Association, 2014, pp. 305–320. [Online]. Available: http://dl.acm.org/citation.cfm?id=2643634.2643666
2014
Cited alongside, same era.
M. Li, D. G. Andersen, A. J. Smola, and K. Yu, “Communication efficient distributed machine learning with the parameter server,” in Advances in Neural Information Processing Systems , 2014, pp. 19–27
2014
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Gibiansky, “Bringing HPC techniques to deep learning,” Baidu Research, Tech. Rep., 2017, http://research. baidu. com/bringing-hpc-techniques-deep-learning/. Bingjing Zhang TESTS & CERTIFICATIONS IBM Certified Database Associate-DB2 Universal Database, Tech. Rep., 2017
2017
Later among the works it cites.
L. N. Smith, “Cyclical learning rates for training neural networks,” in Applications of Computer Vision (WACV), 2017 IEEE Winter Conference on . IEEE, 2017, pp. 464–472
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu, “1-Bit Stochastic Gradient Descent and Application to Data-Parallel Distributed Training of Speech DNNs,” in Interspeech 2014 , September 2014
2014
Cited alongside, same era.
2015
Cited alongside, same era.
S. Zhang, A. E. Choromanska, and Y. LeCun, “Deep learning with elastic averaging SGD,” in Advances in Neural Information Processing Systems , 2015, pp. 685–693
2015
Cited alongside, same era.
N. Strom, “Scalable distributed DNN training using commodity GPU cloud computing,” in INTERSPEECH , 2015
2015
Cited alongside, same era.
2016
Cited alongside, same era.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al. , “Tensorflow: a system for large-scale machine learning.” in OSDI , vol. 16, 2016, pp. 265–283
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
K. Greff, R. K. Srivastava, J. Koutník, B. R. Steunebrink, and J. Schmidhuber, “LSTM: A search space odyssey,” IEEE transactions on neural networks and learning systems , vol. 28, no. 10, pp. 2222–2232, 2017
2017
Later among the works it cites.
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, “Terngrad: Ternary gradients to reduce communication in distributed deep learning,” in Advances in neural information processing systems , 2017, pp. 1509–1519
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2018
Closest in time.
B. Ginsburg, I. Gitman, and Y. You, “Large Batch Training of Convolutional Networks with Layer-wise Adaptive Rate Scaling,” 2018
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.