Fetching the paper…
Reading the bibliography…
Neural transducer-based systems such as RNN Transducers (RNN-T) for automatic speech recognition (ASR) blend the individual components of a traditional hybrid ASR systems (acoustic model, language model, punctuation model, inverse text normalization) into one single model.
“A one-pass decoder based on polymorphic linguistic context assignment,”
Hagen Soltau, Florian Metze, Christian Fugen, and Alex Waibel, · 2001
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Rectified linear units improve restricted boltzmann machines,”
Vinod Nair and Geoffrey E Hinton, · 2010
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition,”
Geoffrey Hinton, Li Deng, Dong Yu, George Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Brian Kingsbury, et al., · 2012
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling,”
Haşim Sak, Andrew Senior, and Françoise Beaufays, · 2014
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
“Highway long short-term memory rnns for distant speech recognition,”
Yu Zhang, Guoguo Chen, Dong Yu, Kaisheng Yaco, Sanjeev Khudanpur, and James Glass, · 2016
Cited alongside, same era.
“Purely sequence-trained neural networks for asr based on lattice-free mmi.,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
“Improving latency-controlled blstm acoustic models for online speech recognition,”
Shaofei Xue and Zhijie Yan, · 2017
Later among the works it cites.
Taku Kudo and John Richardson, · 2018
Later among the works it cites.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Closest in time.
“From senones to chenones: Tied context-dependent graphemes for hybrid speech recognition,”
D. Le, X. Zhang, W. Zheng, C. Fuegen, G. Zweig, and M. L. Seltzer, · 2019
Closest in time.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Cited alongside, same era.
Closest in time.