Fetching the paper…
Reading the bibliography…
The performance of automatic speech recognition (ASR) has improved tremendously due to the application of deep neural networks (DNNs).
“Phoneme recognition using time-delay neural networks,”
Alex Waibel, Toshiyuki Hanazawa, Geoffrey Hinton, Kiyohiro Shikano, and Kevin J Lang, · 1989
Earlier work this paper cites.
“A tutorial on hidden Markov models and selected applications in speech recognition,”
Lawrence R Rabiner, · 1989
Earlier work this paper cites.
“A novel objective function for improved phoneme recognition using time-delay neural networks,”
John B Hampshire, Alexander H Waibel, et al., · 1990
Earlier work this paper cites.
“Learning long-term dependencies with gradient descent is difficult,”
Yoshua Bengio, Patrice Simard, and Paolo Frasconi, · 1994
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“A one-pass decoder based on polymorphic linguistic context assignment,”
Hagen Soltau, Florian Metze, Christian Fügen, and Alex Waibel, · 2001
Earlier work this paper cites.
“Weighted finite-state transducers in speech recognition,”
Mehryar Mohri, Fernando Pereira, and Michael Riley, · 2002
Earlier work this paper cites.
“Learning precise timing with LSTM recurrent networks,”
Felix A Gers, Nicol N Schraudolph, and Jürgen Schmidhuber, · 2003
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“A fast learning algorithm for deep belief nets,”
Geoffrey Hinton, Simon Osindero, and Yee-Whye Teh, · 2006
Earlier work this paper cites.
“OpenFst: A general and efficient weighted finite-state transducer library,”
Cyril Allauzen, Michael Riley, Johan Schalkwyk, Wojciech Skut, and Mehryar Mohri, · 2007
Earlier work this paper cites.
“Feature engineering in context-dependent deep neural networks for conversational speech transcription,”
Frank Seide, Gang Li, Xie Chen, and Dong Yu, · 2011
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukáš Burget, Ondřej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlíček, Yanmin Qian, Petr Schwarz, Jan Silovský, Georg Stemmer, and Karel Veselý, · 2011
Earlier work this paper cites.
“Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
George E Dahl, Dong Yu, Li Deng, and Alex Acero, · 2012
Cited alongside, same era.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al., · 2012
Cited alongside, same era.
“Adaptation of context-dependent deep neural networks for automatic speech recognition,”
Kaisheng Yao, Dong Yu, Frank Seide, Hang Su, Li Deng, and Yifan Gong, · 2012
Cited alongside, same era.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Cited alongside, same era.
“Hybrid speech recognition with deep bidirectional LSTM,”
Alex Graves, Navdeep Jaitly, and Abdel-rahman Mohamed, · 2013
Cited alongside, same era.
“Deep convolutional neural networks for large-scale speech tasks,”
Tara N Sainath, Brian Kingsbury, George Saon, Hagen Soltau, Abdel-rahman Mohamed, George Dahl, and Bhuvana Ramabhadran, · 2014
Later among the works it cites.
“Distributed learning of multilingual DNN feature extractors using GPUs,”
Yajie Miao, Hao Zhang, and Florian Metze, · 2014
Later among the works it cites.
“Improving language-universal feature extraction with deep maxout and convolutional neural networks,”
Yajie Miao and Florian Metze, · 2014
Later among the works it cites.
“Towards speaker adaptive training of deep neural network acoustic models,”
Yajie Miao, Hao Zhang, and Florian Metze, · 2014
Later among the works it cites.
“Lexicon-free conversational speech recognition with neural networks,”
Andrew L Maas, Ziang Xie, Dan Jurafsky, and Andrew Y Ng, · 2015
Closest in time.
“End-to-end attention-based large vocabulary speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Speaker adaptation of context dependent deep neural networks,”
Hank Liao, · 2013
Cited alongside, same era.
“GMM-free DNN training,”
Andrew Senior, Georg Heigold, Michiel Bacchiani, and Hank Liao, · 2014
Cited alongside, same era.
“Asynchronous, online, GMM-free training of a context dependent acoustic model for speech recognition,”
Michiel Bacchiani, Andrew Senior, and Georg Heigold, · 2014
Cited alongside, same era.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Cited alongside, same era.
“Deepspeech: Scaling up end-to-end speech recognition,”
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al., · 2014
Cited alongside, same era.
“First-pass large vocabulary continuous speech recognition using bi-directional recurrent DNNs,”
Awni Y Hannun, Andrew L Maas, Daniel Jurafsky, and Andrew Y Ng, · 2014
Cited alongside, same era.
“End-to-end continuous speech recognition using attention-based recurrent NN: First results,”
Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Cited alongside, same era.
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2015
Closest in time.
William Chan, Navdeep Jaitly, Quoc V Le, and Oriol Vinyals, · 2015
Closest in time.
“Learning acoustic frame labeling for speech recognition with recurrent neural networks,”
Hasim Sak, Andrew Senior, Kanishka Rao, Ozan Irsoy, Alex Graves, Francoise Beaufays, and Johan Schalkwyk, · 2015
Closest in time.
“Convolutional, long short-term memory, fully connected deep neural networks,”
Tara N Sainath, Oriol Vinyals, Andrew Senior, and Hasim Sak, · 2015
Closest in time.
“On speaker adaptation of long short-term memory recurrent neural networks,”
Yajie Miao and Florian Metze, · 2015
Closest in time.
“Towards end-to-end speech recognition for chinese mandarin using long short-term memory recurrent neural networks,”
Jie Li, Heng Zhang, Xinyuan Cai, and Bo Xu, · 2015
Closest in time.
“Speaker adaptive training of deep neural network acoustic models using i-vectors,”
Yajie Miao, Hao Zhang, and Florian Metze, · 2015
Closest in time.