Fetching the paper…
Reading the bibliography…
The recurrent neural network transducer (RNN-T) has recently become the mainstream end-to-end approach for streaming automatic speech recognition (ASR).
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
Statistical Methods for Speech Recognition
F. Jelinek, · 1997
Earlier work this paper cites.
“Separating style and content with bilinear models,”
J. Tenenbaum and W. Freeman, · 2000
Earlier work this paper cites.
“Connectionist temporal classification: Labeling unsegmented sequenece data with recurrent neural networks,”
A. Graves, S. Fernandez, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Bilinear classifiers for visual recognition,”
H. Pirsiavash, D. Ramanan, and C. Fowlkes, · 2009
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Developing real-time streaming Transformer transducer for speech recognition on large-scale dataset,”
X. Chen, Y. Wu, Z. Wang, S. Liu, and J. Li, · 2012
Earlier work this paper cites.
“Japanese and Korean voice search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
A. Graves, A. Mohamed, and G. Hinton, · 2013
Earlier work this paper cites.
“Attention-based models for speech recognition,”
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“A study of the recurrent neural network encoder-decoder for large vocabulary speech recognition,”
L. Lu, X. Zhang, K. Cho, and S. Renals, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q.V. Le, and O. Vinyals, · 2016
Earlier work this paper cites.
“Highway long short-term memory RNNs for distant speech recognition,”
Y. Zhang, G. Chen, D. Yu, K. Yao, S. Khudanpur, and J. Glass, · 2016
Earlier work this paper cites.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
S. Kim, T. Hori, and S. Watanabe, · 2017
Earlier work this paper cites.
“A comparison of sequence-to-sequence models for speech recognition,”
R. Prabhavalkar, K. Rao, T. Sainath, B. Li, L. Johnson, and N. Jaitly, · 2017
Earlier work this paper cites.
“Exploring neural transducers for end-to-end speech recognition,”
E. Battenberg, J. Chen, R. Child, A. Coates, Y. Gaur, Y. Li, H. Liu, S. Satheesh, D. Seetapun, A. Sriram, and Z. Zhu, · 2017
Cited alongside, same era.
“Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,”
K. Rao, H. Sak, and R. Prabhavalkar, · 2017
Cited alongside, same era.
“Hadamard product for low-rank bilinear pooling,”
J.H. Kim, K.W. On, W. Lim, J. Kim, J.W. Ha, and B.T. Zhang, · 2017
Cited alongside, same era.
“Overcoming language priors in visual question answering with adversarial regularization,”
S. Ramakrishnan, A. Agrawal, , and S. Lee, · 2018
Cited alongside, same era.
“Improving RNN transducer modeling for end-to-end speech recognition,”
J. Li, R. Zhao, H. Hu, and Y. Gong, · 2019
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
“Hybrid autoregressive transducer (HAT),”
E. Variani, D. Rybach, C. Allauzen, and M. Riley, · 2020
Later among the works it cites.
“Developing RNN-T models surpassing high-performance hybrid models with customization capability,”
J. Li, R. Zhao, Z. Meng, Y. Liu, W. Wei, S. Parthasarathy, V. Mazalov, Z. Wang, L. He, S. Zhao, and Y. Gong, · 2020
Later among the works it cites.
“Multimodal intelligence: Representation learning, information fusion, and applications,”
C. Zhang, Z. Yang, X. He, and L. Deng, · 2020
Later among the works it cites.
“SpecAugment on large scale datasets,”
D.S. Park, Y. Zhang, C.-C. Chiu, Y. Chen, B. Li, W. Chan, Q.V. Le, and Y. Wu, · 2020
Later among the works it cites.
“Contextualized streaming end-to-end speech recognition with Trie-based deep biasing and shallow fusion,”
D. Le, M. Jain, G. Keren, S. Kim, Y. Shi, J. Mahadeokar, J. Chan, Y. Shangguan, C. Fuegen, O. Kalinli, Y. Saraf, and M.L. Seltzer, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. He, T. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang, Q. Liang, D. Bhatia, Y. Shangguan, B. Li, G. Pundak, K.C. Sim, T. Bagby, S. Chang, R. Rao, and A. Gruenstein, · 2019
Cited alongside, same era.
“Joint endpointing and decoding with end-to-end models,”
S. Chang, R. Prabhavalkar, Y. He, T.N. Sainath, and G. Simko, · 2019
Cited alongside, same era.
“A density ratio approach to language model fusion in end-to-end automatic speech recognition,”
E. McDermott, H. Sak, and E. Variani, · 2019
Cited alongside, same era.
“Lingvo: A modular and scalable framework for sequence-to-sequence modeling,”
J. Shen, P. Nguyen, et al., · 2019
Cited alongside, same era.
“A new training pipeline for an improved neural transducer,”
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, · 2020
Cited alongside, same era.
“Towards fast and accurate streaming end-to-end ASR,”
B. Li, S. Chang, T.N. Sainath, R. Pang, Y. He, T. Strohman, and Y. Wu, · 2020
Cited alongside, same era.
“Transformer Transducer: A streamable speech recognition model with Transformer encoders and RNN-T loss,”
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, · 2020
Cited alongside, same era.
“Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,”
Y. Shi, Y. Wang, C. Wu, C.-F. Yeh, J. Chan, F. Zhang, D. Le, and M. Seltzer, · 2021
Later among the works it cites.
“Cascaded encoders for unifying streaming and non-streaming ASR,”
A. Narayanan, T.N. Sainath, R. Pang, J. Yu, C.-C. Chiu, R. Prabhavalkar, E. Variani, and T. Strohman, · 2021
Later among the works it cites.
“Tiny transducer: A highly-efficient speech recognition model on edge devices,”
Y. Zhang, S. Sun, and L. Ma, · 2021
Later among the works it cites.
“Tree-constrained pointer generator for end-to-end contextual speech recognition,”
G. Sun, C. Zhang, and P.C. Woodland, · 2021
Later among the works it cites.
“Internal language model training for domain-adaptive end-to-end speech recognition,”
Z. Meng, N. Kanda, Y. Gaur, S. Parthasarathy, E. Sun, L. Lu, X. Chen, J. Li, and Y. Gong, · 2021
Later among the works it cites.
“Using synthetic audio to improve the recognition of out-of-vocabulary words in end-to-end ASR systems,”
X. Zheng, Y. Liu, D. Gunceler, and D. Willett, · 2021
Later among the works it cites.
“Improving streaming automatic speech recognition with non-streaming model distillation on unsupervised data,”
T. Doutre, W. Han, M. Ma, Z. Lu, C.-C. Chiu, R. Pang, A. Narayanan, A. Misra, Y. Zhang, and L. Cao, · 2021
Later among the works it cites.
“Advancing RNN transducer technology for speech recognition,”
G. Saon, Z. Tüske, D. Bolanos, and B. Kingsbury, · 2021
Later among the works it cites.
“Scaling end-to-end models for large-scale multilingual ASR,”
B. Li, R. Pang, T.N. Sainath, A. Gulati, Y. Zhang, J. Qin, P. Haghani, W.R. Huang, and M. Ma, · 2021
Later among the works it cites.
“Combination of deep speaker embeddings for diarisation,”
G. Sun, C. Zhang, and P.C. Woodland, · 2021
Later among the works it cites.