Fetching the paper…
Reading the bibliography…
We investigate a set of techniques for RNN Transducers (RNN-Ts) that were instrumental in lowering the word error rate on three different tasks (Switchboard 300 hours, conversational Spanish 780 hours and conversational Italian 900 hours).
“The Kaldi speech recognition toolkit,”
D. Povey, A. Ghoshal, G. Boulianne, et al., · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
A. Graves, A.-r. Mohamed, and G. Hinton, · 2013
Earlier work this paper cites.
“Regularization of neural networks using DropConnect,”
L. Wan, M. Zeiler, S. Zhang, et al., · 2013
Earlier work this paper cites.
“Speaker adaptation of neural network acoustic models using i-vectors,”
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, · 2013
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, et al., · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Earlier work this paper cites.
Y. Wu, M. Schuster, Z. Chen, et al., · 2016
Earlier work this paper cites.
“On multiplicative integration with recurrent neural networks,”
Y. Wu, S. Zhang, Y. Zhang, et al., · 2016
Earlier work this paper cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,”
K. Rao, H. Sak, and R. Prabhavalkar, · 2017
Earlier work this paper cites.
“Exploring neural transducers for end-to-end speech recognition,”
E. Battenberg, J. Chen, R. Child, et al., · 2017
Earlier work this paper cites.
“Exploring RNN-Transducer for Chinese speech recognition,”
S. Wang, P. Zhou, W. Chen, et al., · 2018
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
C.-C. Chiu, T. N. Sainath, Y. Wu, et al., · 2018
Cited alongside, same era.
“Switchout: an efficient data augmentation algorithm for neural machine translation,”
X. Wang, H. Pham, Z. Dai, and G. Neubig, · 2018
Cited alongside, same era.
“Efficient implementation of recurrent neural network transducer in TensorFlow,”
T. Bagby, K. Rao, and K. C. Sim, · 2018
Cited alongside, same era.
“End-to-end speech recognition using lattice-free MMI,”
H. Hadian, H. Sameti, D. Povey, and S. Khudanpur, · 2018
Cited alongside, same era.
“Multiplicative interactions and where to find them,”
S. M. Jayakumar, W. M. Czarnecki, J. Menick, et al., · 2019
Later among the works it cites.
“Forget a bit to learn better: Soft forgetting for CTC-based automatic speech recognition.,”
K. Audhkhasi, G. Saon, Z. Tüske, et al., · 2019
Later among the works it cites.
“Super-convergence: Very fast training of neural networks using large learning rates,”
L. N. Smith and N. Topin, · 2019
Later among the works it cites.
“Training language models for long-span cross-sentence evaluation,”
K. Irie, A. Zeyer, R. Schlüter, and H. Ney, · 2019
Later among the works it cites.
“Transformer-transducer: End-to-end speech recognition with self-attention,”
C.-F. Yeh, J. Mahadeokar, K. Kalgaonkar, et al., · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Streaming end-to-end speech recognition for mobile devices,”
Y. He, T. N. Sainath, R. Prabhavalkar, et al., · 2019
Cited alongside, same era.
“Improving RNN transducer modeling for end-to-end speech recognition,”
J. Li, R. Zhao, H. Hu, and Y. Gong, · 2019
Cited alongside, same era.
“Monotonic recurrent neural network transducer and decoding strategies,”
A. Tripathi, H. Lu, H. Sak, and H. Soltau, · 2019
Cited alongside, same era.
“RNN-T for latency controlled ASR with improved beam search,”
M. Jain, K. Schubert, J. Mahadeokar, et al., · 2019
Cited alongside, same era.
“A density ratio approach to language model fusion in end-to-end automatic speech recognition,”
E. McDermott, H. Sak, and E. Variani, · 2019
Cited alongside, same era.
“Sequence noise injected training for end-to-end speech recognition,”
G. Saon, Z. Tüske, K. Audhkhasi, and B. Kingsbury, · 2019
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, Y. Zhang, et al., · 2019
Cited alongside, same era.
C. Weng, C. Yu, J. Cui, et al., · 2019
Later among the works it cites.
“Hybrid autoregressive transducer (HAT),”
E. Variani, D. Rybach, C. Allauzen, and M. Riley, · 2020
Later among the works it cites.
“A new training pipeline for an improved neural transducer,”
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, · 2020
Later among the works it cites.
“RNN-T models fail to generalize to out-of-domain audio: Causes and solutions,”
C.-C. Chiu, A. Narayanan, W. Han, et al., · 2020
Later among the works it cites.
“Alignment-length synchronous decoding for RNN transducer,”
G. Saon, Z. Tüske, and K. Audhkhasi, · 2020
Later among the works it cites.
“Single headed attention based sequence-to-sequence model for state-of-the-art results on Switchboard-300,”
Z. Tüske, G. Saon, K. Audhkhasi, and B. Kingsbury, · 2020
Later among the works it cites.
“Knowledge distillation from offline to streaming RNN transducer for end-to-end speech recognition,”
G. Kurata and G. Saon, · 2020
Later among the works it cites.