H. Robbins and D. Siegmund, “A convergence theorem for non negative almost supermartingales and some applications,” in Herbert Robbins Selected Papers . Springer, 1985, pp. 111–135
1985
Earlier work this paper cites.
W. Gropp, E. Lusk, N. Doss, and A. Skjellum, “A High-Performance, Portable Implementation of the MPI Message Passing Interface Standard,” Parallel Computing , vol. 22, no. 6, pp. 789–828, 1996
1996
Earlier work this paper cites.
A. Geist, W. Gropp, S. Huss-Lederman, A. Lumsdaine, E. L. Lusk, W. Saphir, T. Skjellum, and M. Snir, “MPI-2: Extending the message-passing interface,” in Euro-Par, Vol. I , 1996, pp. 128–135
1996
Earlier work this paper cites.
D. Saad, “Online algorithms and stochastic approximations,” Online Learning , vol. 5, 1998
1998
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
K. C. Kiwiel, “Convergence and efficiency of subgradient methods for quasiconvex minimization,” Mathematical programming , vol. 90, no. 1, pp. 1–25, 2001
2001
Earlier work this paper cites.
R. Collobert, S. Bengio, and J. Marithoz, “Torch: A modular machine learning software library,” 2002
2002
Earlier work this paper cites.
S. Sur, H.-W. Jin, L. Chai, and D. K. Panda, “Rdma read based rendezvous protocol for mpi over infiniband: design alternatives and benefits.” in PPOPP , 2006, pp. 32–39
2006
Earlier work this paper cites.
G. E. Hinton and S. Osindero, “A fast learning algorithm for deep belief nets,” Neural Computation , vol. 18, p. 2006, 2006
2006
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” Science , vol. 313, no. 5786, pp. 504–507, Jul. 2006. [Online]. Available: http://www.ncbi.nlm.nih.gov/sites/entrez?db=pubmed&uid=16873662&cmd=showdetailview&indexed=google
2006
Earlier work this paper cites.
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, “Greedy layer-wise training of deep networks,” in Advances in Neural Information Processing Systems 19 , B. Schölkopf, J. C. Platt, and T. Hoffman, Eds. MIT Press, 2007, pp. 153–160. [Online]. Available: http://papers.nips.cc/paper/3048-greedy-layer-wise-training-of-deep-networks.pdf
2007
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” University of Toronto, Tech. Rep., 2009
2009
Earlier work this paper cites.
J. Bergstra, O. Breuleux, F. Bastien, P. Lamblin, R. Pascanu, G. Desjardins, J. Turian, D. Warde-Farley, and Y. Bengio, “Theano: a CPU and GPU math expression compiler,” in Proceedings of the Python for Scientific Computing Conference (SciPy) , Jun. 2010, oral Presentation
2010
Earlier work this paper cites.
T. Hoefler, T. Schneider, and A. Lumsdaine, “Characterizing the Influence of System Noise on Large-Scale Applications by Simulation,” in International Conference for High Performance Computing, Networking, Storage and Analysis (SC’10) , Nov. 2010
2010
Earlier work this paper cites.
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.-A. Manzagol, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J. Mach. Learn. Res. , vol. 11, pp. 3371–3408, Dec. 2010. [Online]. Available: http://dl.acm.org/citation.cfm?id=1756006.1953039
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25 , F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds. Curran Associates, Inc., 2012, pp. 1097–1105. [Online]. Available: http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf
2012
Earlier work this paper cites.
F. Bastien, P. Lamblin, R. Pascanu, J. Bergstra, I. J. Goodfellow, A. Bergeron, N. Bouchard, and Y. Bengio, “Theano: new features and speed improvements,” Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop, 2012
2012
Earlier work this paper cites.
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng, “Large scale distributed deep networks,” in Proceedings of the 25th International Conference on Neural Information Processing Systems , ser. NIPS’12. USA: Curran Associates Inc., 2012, pp. 1223–1231. [Online]. Available: http://dl.acm.org/citation.cfm?id=2999134.2999271
2012
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng, “Large scale distributed deep networks,” in Advances in Neural Information Processing Systems 25 , P. Bartlett, F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds., 2012, pp. 1232–1240. [Online]. Available: http://books.nips.cc/papers/files/nips25/NIPS2012_0598.pdf
2012
Earlier work this paper cites.
Y. LeCun, C. Cortes, and C. J. Burges, “The mnist database of handwritten digits, 1998,” Available electronically at http://yann.lecun.com/exdb/mnist , 2012
2012
Earlier work this paper cites.
T. Tieleman and G. Hinton, “Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural networks for machine learning , vol. 4, no. 2, pp. 26–31, 2012
2012
Earlier work this paper cites.
A. Bhatele, K. Mohror, S. H. Langer, and K. E. Isaacs, “There goes the neighborhood: Performance degradation due to nearby jobs,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis , ser. SC ’13. New York, NY, USA: ACM, 2013, pp. 41:1–41:12. [Online]. Available: http://doi.acm.org/10.1145/2503210.2503247
2013
Earlier work this paper cites.
Q. Ho, J. Cipar, H. Cui, J. K. Kim, S. Lee, P. B. Gibbons, G. A. Gibson, G. R. Ganger, and E. P. Xing, “More effective distributed ml via a stale synchronous parallel parameter server,” in Proceedings of the 26th International Conference on Neural Information Processing Systems , ser. NIPS’13. USA: Curran Associates Inc., 2013, pp. 1223–1231. [Online]. Available: http://dl.acm.org/citation.cfm?id=2999611.2999748
2013
Earlier work this paper cites.