Fetching the paper…
Reading the bibliography…
This paper addresses the observed performance gap between automatic speech recognition (ASR) systems based on Long Short Term Memory (LSTM) neural networks trained with the connectionist temporal classification (CTC) loss function and systems based on hybrid Deep Neural Networks (DNNs) trained with the cross entropy (CE) loss function on domains with limited data.
Optimization by simulated annealing
S. Kirkpatrick, C. D. Gelatt, and M. P. Vecchi · 1983
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K. Paliwal · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Felix A. Gers, Jürgen Schmidhuber, and Fred A. Cummins · 2000
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely · 2011
Earlier work this paper cites.
Vocal tract length perturbation (vtlp) improves speech recognition
Navdeep Jaitly and G. E. Hinton · 2013
Earlier work this paper cites.
Sequence-discriminative training of deep neural networks
Karel Veselý, Arnab Ghoshal, Lukás Burget, and Daniel Povey · 2013
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
Alex Graves and Navdeep Jaitly · 2014
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Awni Y. Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, and Andrew Y. Ng · 2014
Earlier work this paper cites.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling
Hasim Sak, Andrew W. Senior, and Françoise Beaufays · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Dropout improves recurrent neural networks for handwriting recognition
Vu Pham, Théodore Bluche, Christopher Kermorvant, and Jérôme Louradour · 2014
Cited alongside, same era.
Fast and accurate recurrent neural network acoustic models for speech recognition
Hasim Sak, Andrew W. Senior, Kanishka Rao, and Françoise Beaufays · 2015
Cited alongside, same era.
EESEN: end-to-end speech recognition using deep RNN models and wfst-based decoding
Yajie Miao, Mohammad Gowayyed, and Florian Metze · 2015
Cited alongside, same era.
A time delay neural network architecture for efficient modeling of long temporal contexts
Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur · 2015
Later among the works it cites.
Convolutional, long short-term memory, fully connected deep neural networks
Tara N. Sainath, Oriol Vinyals, Andrew W. Senior, and Hasim Sak · 2015
Later among the works it cites.
Acoustic modelling with CD-CTC-SMBR LSTM RNNS
Andrew W. Senior, Hasim Sak, Felix de Chaumont Quitry, Tara N. Sainath, and Kanishka Rao · 2015
Later among the works it cites.
Deep speech 2 : End-to-end speech recognition in english and mandarin
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, Erich Elsen, Jesse Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang, Yi Wang, Zhiqian Wang, Bo Xiao, Yan Xie, Dani Yogatama, Jun Zhan, and Zhenyao Zhu · 2016
Later among the works it cites.
Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Cited alongside, same era.
An empirical exploration of recurrent network architectures
Rafal Józefowicz, Wojciech Zaremba, and Ilya Sutskever · 2015
Cited alongside, same era.
Audio augmentation for speech recognition
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur · 2015
Cited alongside, same era.
Data augmentation for deep convolutional neural network acoustic modeling
Xiaodong Cui, Vaibhava Goel, and Brian Kingsbury · 2015
Cited alongside, same era.
RNNDROP: A novel dropout for RNNS in ASR
Taesup Moon, Heeyoul Choi, Hoshik Lee, and Inchul Song · 2015
Cited alongside, same era.
Hagen Soltau, Hank Liao, and Hasim Sak · 2016
Later among the works it cites.
Purely sequence-trained neural networks for ASR based on lattice-free MMI
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur · 2016
Later among the works it cites.
An empirical exploration of CTC acoustic models
Yajie Miao, Mohammad Gowayyed, Xingyu Na, Tom Ko, Florian Metze, and Alexander H. Waibel · 2016
Later among the works it cites.
Recurrent dropout without memory loss
Stanislau Semeniuta, Aliaksei Severyn, and Erhardt Barth · 2016
Later among the works it cites.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Later among the works it cites.
An exploration of dropout with lstms
Gaofeng Cheng, Vijayaditya Peddinti, Daniel Povey, Vimal Manohar, Sanjeev Khudanpur, and Yonghong Yan · 2017
Closest in time.