Fetching the paper…
Reading the bibliography…
A simplified speech recognition system that uses the maximum mutual information (MMI) criterion is considered.
“Maximum mutual information estimation of hidden markov model parameters for speech recognition,”
Lalit Bahl, Peter Brown, Peter De Souza, and Robert Mercer, · 1986
Earlier work this paper cites.
“A tutorial on hidden Markov models and selected applications in speech recognition,”
R. L. Rabiner, · 1989
Earlier work this paper cites.
“The design for the Wall Street Journal-based CSR corpus,”
Douglas B Paul and Janet M Baker, · 1992
Earlier work this paper cites.
Fundamentals of Speech Recognition
Lawrence R Rabiner and Biing-Hwang Juang, · 1993
Earlier work this paper cites.
“A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (rover),”
Jonathan G Fiscus, · 1997
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Weighted finite-state transducers in speech recognition,”
Mehryar Mohri, Fernando Pereira, and Michael Riley, · 2002
Earlier work this paper cites.
“Learning precise timing with LSTM recurrent networks,”
Felix A Gers, Nicol N Schraudolph, and Jürgen Schmidhuber, · 2002
Earlier work this paper cites.
Discriminative training for large vocabulary speech recognition
Daniel Povey, · 2005
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“OpenFst: A general and efficient weighted finite-state transducer library. Implementation and Application of Automata,” 2007
Cyril Allauzen, Michael Riley, Johan Schalkwyk, Wojciech Skut, and Mehryar Mohri, · 2007
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, et al., · 2011
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
Geoffrey Hinton, Li Deng, Dong Yu, et al., · 2012
Earlier work this paper cites.
“Sequence-discriminative training of deep neural networks.,”
Karel Veselỳ, Arnab Ghoshal, Lukás Burget, and Daniel Povey, · 2013
Earlier work this paper cites.
“On the difficulty of training recurrent neural networks,”
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio, · 2013
Cited alongside, same era.
“Towards end-to-end speech recognition with recurrent neural networks.,”
Alex Graves and Navdeep Jaitly, · 2014
Cited alongside, same era.
“Deep speech: Scaling up end-to-end speech recognition,”
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al., · 2014
Cited alongside, same era.
“Deep speech: scaling up end-to-end speech recognition,”
Awni Y. Hannun, Carl Case, Jared Casper, Bryan C. Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, and Andrew Y. Ng, · 2014
Cited alongside, same era.
“Ensemble deep learning for speech recognition,”
Li Deng and John Platt, · 2014
Cited alongside, same era.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Later among the works it cites.
“Latent sequence decompositions,”
William Chan, Yu Zhang, Quoc Le, and Navdeep Jaitly, · 2016
Later among the works it cites.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2016
Later among the works it cites.
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al., · 2016
Later among the works it cites.
“Professor forcing: A new algorithm for training recurrent networks,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Kaldi + PDNN: building DNN-based ASR systems with kaldi and PDNN,”
Yajie Miao, · 2014
Cited alongside, same era.
“ADAM: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Deep recurrent neural networks for acoustic modelling,”
William Chan and Ian Lane, · 2015
Cited alongside, same era.
William Chan, Navdeep Jaitly, Quoc V Le, and Oriol Vinyals, · 2015
Cited alongside, same era.
“Learning acoustic frame labeling for speech recognition with recurrent neural networks,”
Haşim Sak, Andrew Senior, Kanishka Rao, Ozan Irsoy, Alex Graves, Françoise Beaufays, and Johan Schalkwyk, · 2015
Cited alongside, same era.
“EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Cited alongside, same era.
“Acoustic modelling with CD-CTC-sMBR LSTM RNNs,”
Haşim Sak, Félix de Chaumont Quitry, Tara Sainath, Kanishka Rao, et al., · 2015
Cited alongside, same era.
Alex M Lamb, Anirudh Goyal ALIAS PARTH GOYAL, Ying Zhang, Saizheng Zhang, Aaron C Courville, and Yoshua Bengio, · 2016
Later among the works it cites.
“Maximum a posteriori based decoding for CTC acoustic models,”
Naoyuki Kanda, Xugang Lu, and Hisashi Kawai, · 2016
Later among the works it cites.
“Purely sequence-trained neural networks for ASR based on lattice-free MMI.,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Later among the works it cites.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,”
Martın Abadi, Ashish Agarwal, et al., · 2016
Later among the works it cites.
“Kaldi with TensorFlow Neural Net,” https://github.com/vrenkens/tfkaldi , 2016
Vincent Renkens, · 2016
Later among the works it cites.
“Personalized speech recognition on mobile devices,”
Ian McGraw, Rohit Prabhavalkar, Raziel Alvarez, et al., · 2016
Later among the works it cites.
“Residual convolutional CTC networks for automatic speech recognition,”
Yisen Wang, Xuejiao Deng, Songbai Pu, and Zhiheng Huang, · 2017
Closest in time.
“English conversational telephone speech recognition by humans and machines,”
George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, , et al., · 2017
Closest in time.
Cheng Ju, Aurélien Bibaut, and Mark J van der Laan, · 2017
Closest in time.