Fetching the paper…
Reading the bibliography…
The creation of practical deep learning data-products often requires parallelization across processors and computers to make deep learning feasible on large data sets, but bottlenecks in communication bandwidth make it difficult to attain good speedups through parallelism.
Learning representations by back-propagating errors
Rumelhart, David E, Hinton, Geoffrey E, and Williams, Ronald J · 1988
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
Hochreiter, Sepp, Bengio, Yoshua, Frasconi, Paolo, and Schmidhuber, Jürgen · 2001
Earlier work this paper cites.
High performance rdma based all-to-all broadcast for infiniband clusters
Sur, Sayantan, Bondhugula, Uday Kumar Reddy, Mamidala, Amith, Jin, H-W, and Panda, Dhabaleswar K · 2005
Earlier work this paper cites.
Improving the speed of neural networks on cpus
Vanhoucke, Vincent, Senior, Andrew, and Mao, Mark Z · 2011
Earlier work this paper cites.
Multi-column deep neural networks for image classification
Ciresan, Dan, Meier, Ueli, and Schmidhuber, Jürgen · 2012
Earlier work this paper cites.
Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition
Dahl, George E, Yu, Dong, Deng, Li, and Acero, Alex · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Dean, Jeffrey, Corrado, Greg, Monga, Rajat, Chen, Kai, Devin, Matthieu, Mao, Mark, Senior, Andrew, Tucker, Paul, Yang, Ke, Le, Quoc V, et al · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
Optimizing All-to-All and Allgather Communications on GPGPU Clusters
Singh, Ashish Kumar · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Cited alongside, same era.
Deep learning with cots hpc systems
Coates, Adam, Huval, Brody, Wang, Tao, Wu, David, Catanzaro, Bryan, and Andrew, Ng · 2013
Cited alongside, same era.
Low precision arithmetic for deep learning
Courbariaux, Matthieu, Bengio, Yoshua, and David, Jean-Pierre · 2014
Later among the works it cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, Alex · 2014
Later among the works it cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Seide, Frank, Fu, Hao, Droppo, Jasha, Li, Gang, and Yu, Dong · 2014
Later among the works it cites.
Deep learning with limited numerical precision
Gupta, Suyog, Agrawal, Ankur, Gopalakrishnan, Kailash, and Narayanan, Pritish · 2015
Closest in time.
Deep learning in neural networks: An overview
Schmidhuber, Jürgen · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Goodfellow, Ian J, Warde-Farley, David, Mirza, Mehdi, Courville, Aaron, and Bengio, Yoshua · 2013
Cited alongside, same era.
Project adam: Building an efficient and scalable deep learning training system
Chilimbi, Trishul, Suzue, Yutaka, Apacible, Johnson, and Kalyanaraman, Karthik · 2014
Cited alongside, same era.
Scalable distributed dnn training using commodity gpu cloud computing
Strom, Nikko · 2015
Closest in time.
Deep image: Scaling up image recognition
Wu, Ren, Yan, Shengen, Shan, Yi, Dang, Qingqing, and Sun, Gang · 2015
Closest in time.