Fetching the paper…
Reading the bibliography…
Recurrent neural networks (RNNs), especially long short-term memory (LSTM) RNNs, are effective network for sequential task like speech recognition.
Hidden Markov models, maximum mutual information estimation, and the speech recognition problem
Yves Normandin, · 1991
Earlier work this paper cites.
“Acceleration of stochastic approximation by averaging,”
Boris T Polyak and Anatoli B Juditsky, · 1992
Earlier work this paper cites.
Discriminative training for large vocabulary speech recognition
Daniel Povey, · 2005
Earlier work this paper cites.
“Reducing the dimensionality of data with neural networks,”
Geoffrey E Hinton and Ruslan R Salakhutdinov, · 2006
Earlier work this paper cites.
“A fast learning algorithm for deep belief nets,”
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh, · 2006
Earlier work this paper cites.
“Lattice-based optimization of sequence classification criteria for neural-network acoustic modeling,”
Brian Kingsbury, · 2009
Earlier work this paper cites.
“Understanding the difficulty of training deep feedforward neural networks.,”
Xavier Glorot and Yoshua Bengio, · 2010
Earlier work this paper cites.
“Distributed training strategies for the structured perceptron,”
Ryan McDonald, Keith Hall, and Gideon Mann, · 2010
Earlier work this paper cites.
“Parallelized stochastic gradient descent,”
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola, · 2010
Earlier work this paper cites.
“Towards optimal one pass large scale learning with averaged stochastic gradient descent,”
Wei Xu, · 2011
Cited alongside, same era.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al., · 2012
Cited alongside, same era.
“Large scale distributed deep networks,”
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al., · 2012
Cited alongside, same era.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Cited alongside, same era.
“Hybrid speech recognition with deep bidirectional lstm,”
Alex Graves, Navdeep Jaitly, and Abdel-rahman Mohamed, · 2013
Cited alongside, same era.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2015
Later among the works it cites.
“Fast and accurate recurrent neural network acoustic models for speech recognition,”
Haşim Sak, Andrew Senior, Kanishka Rao, and Françoise Beaufays, · 2015
Later among the works it cites.
“Distilling the knowledge in a neural network,”
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, · 2015
Later among the works it cites.
“Learning acoustic frame labeling for speech recognition with recurrent neural networks,”
Haşim Sak, Andrew Senior, Kanishka Rao, Ozan Irsoy, Alex Graves, Françoise Beaufays, and Johan Schalkwyk, · 2015
Later among the works it cites.
“Acoustic modelling with cd-ctc-smbr lstm rnns,”
Haşim Sak, Félix de Chaumont Quitry, Tara Sainath, Kanishka Rao, et al., · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Asynchronous stochastic gradient descent for dnn training,”
Shanshan Zhang, Ce Zhang, Zhao You, Rong Zheng, and Bo Xu, · 2013
Cited alongside, same era.
“Training and analysing deep recurrent neural networks,”
Michiel Hermans and Benjamin Schrauwen, · 2013
Cited alongside, same era.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling.,”
Hasim Sak, Andrew W Senior, and Françoise Beaufays, · 2014
Cited alongside, same era.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2015
Cited alongside, same era.
Later among the works it cites.
“Understanding intermediate layers using linear classifier probes,”
Guillaume Alain and Yoshua Bengio, · 2016
Later among the works it cites.
“Revisiting distributed synchronous sgd,”
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz, · 2016
Later among the works it cites.
“Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering,”
Kai Chen and Qiang Huo, · 2016
Later among the works it cites.
“Exponential moving average model in parallel speech recognition training,”
Tian Xu, Zhang Jun, Ma Zejun, He Yi, and Wei Juan, · 2017
Closest in time.