Fetching the paper…
Reading the bibliography…
As training data rapid growth, large-scale parallel training with multi-GPUs cluster is widely applied in the neural network model learning currently.We present a new approach that applies exponential moving average method in large-scale parallel training of neural network model.
“Acceleration of stochastic approximation by averaging,”
Boris T Polyak and Anatoli B Juditsky, · 1992
Earlier work this paper cites.
“Distributed training strategies for the structured perceptron,”
Ryan McDonald, Keith Hall, and Gideon Mann, · 2010
Earlier work this paper cites.
“Parallelized stochastic gradient descent,”
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola, · 2010
Earlier work this paper cites.
“Towards optimal one pass large scale learning with averaged stochastic gradient descent,”
Wei Xu, · 2011
Earlier work this paper cites.
“Imagenet classification with deep convolutional neural networks,”
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, · 2012
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al., · 2012
Earlier work this paper cites.
“Large scale distributed deep networks,”
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al., · 2012
Earlier work this paper cites.
“Hybrid speech recognition with deep bidirectional lstm,”
Alex Graves, Navdeep Jaitly, and Abdel-rahman Mohamed, · 2013
Cited alongside, same era.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Cited alongside, same era.
“Asynchronous stochastic gradient descent for dnn training,”
Shanshan Zhang, Ce Zhang, Zhao You, Rong Zheng, and Bo Xu, · 2013
Cited alongside, same era.
“Training and analysing deep recurrent neural networks,”
Michiel Hermans and Benjamin Schrauwen, · 2013
Cited alongside, same era.
“Towards end-to-end speech recognition with recurrent neural networks.,”
Alex Graves and Navdeep Jaitly, · 2014
Cited alongside, same era.
“Neural machine translation by jointly learning to align and translate,”
“Deep speech 2: End-to-end speech recognition in english and mandarin,”
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, et al., · 2015
Later among the works it cites.
“Acoustic modelling with cd-ctc-smbr lstm rnns,”
Haşim Sak, Félix de Chaumont Quitry, Tara Sainath, Kanishka Rao, et al., · 2015
Later among the works it cites.
“Fast and accurate recurrent neural network acoustic models for speech recognition,”
Haşim Sak, Andrew Senior, Kanishka Rao, and Françoise Beaufays, · 2015
Later among the works it cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Later among the works it cites.
“Highway long short-term memory rnns for distant speech recognition,”
Yu Zhang, Guoguo Chen, Dong Yu, Kaisheng Yaco, Sanjeev Khudanpur, and James Glass, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Cited alongside, same era.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling.,”
Hasim Sak, Andrew W Senior, and Françoise Beaufays, · 2014
Cited alongside, same era.
“Revisiting distributed synchronous sgd,”
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz, · 2016
Later among the works it cites.
“Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering,”
Kai Chen and Qiang Huo, · 2016
Later among the works it cites.