Fetching the paper…
Reading the bibliography…
In this paper, we propose a dynamic cascaded encoder Automatic Speech Recognition (ASR) model, which unifies models for different deployment scenarios.
2012
Earlier work this paper cites.
H. Liao, E. McDermott, and A. Senior, “Large Scale Deep Neural Network Acoustic Modeling with Semi-supervised Training Data for YouTube Video Transcription,” in Proc. ASRU , 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in ICONIP , 2015, pp. 577–585
2015
Earlier work this paper cites.
G. Pundak and T. N. Sainath, “Lower frame rate neural network acoustic models,” in Proc. Interspeech , 2016
2016
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in Proc. ICASSP , 2017
2017
Earlier work this paper cites.
C. Kim, A. Misra, K. Chin et al. , “Generation of Large-Scale Simulated Utterances in Virtual Rooms to Train Deep-Neural Networks for Far-Field Speech Recognition in Google Home,” in Proc. Interspeech , 2017
2017
Earlier work this paper cites.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5884–5888
2018
Earlier work this paper cites.
C.-C. Chiu, T. N. Sainath, Y. Wu et al. , “State-of-the-art Speech Recognition With Sequence-to-Sequence Models,” in Proc. ICASSP , 2018
2018
Earlier work this paper cites.
R. Prabhavalkar, T. N. Sainath, Y. Wu, P. Nguyen, Z. Chen, C.-C. Chiu, and A. Kannan, “Minimum word error rate training for attention-based sequence-to-sequence models,” in ICASSP , 2018
2018
Earlier work this paper cites.
D. Wang, X. Wang, and S. Lv, “An overview of end-to-end automatic speech recognition,” Symmetry , vol. 11, no. 8, p. 1018, 2019
2019
Cited alongside, same era.
Y. He, T. N. Sainath, R. Prabhavalkar et al. , “Streaming End-to-end Speech Recognition For Mobile Devices,” in Proc. ICASSP , 2019
2019
Cited alongside, same era.
J. Li, R. Zhao, H. Hu, and Y. Gong, “Improving RNN transducer modeling for end-to-end speech recognition,” in Proc. ASRU , 2019
2019
Cited alongside, same era.
A. Narayanan, R. Prabhavalkar, C.-C. Chiu, D. Rybach, T. N. Sainath, and T. Strohman, “Recognizing long-form speech using streaming end-to-end models,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 920–927
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. Cubuk, and Q. Le, “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” in Proc. Interspeech , 2019
T. N. Sainath, Y. He, A. Narayanan, R. Botros et al. , “An Efficient Streaming Non-Recurrent On-Device End-to-End Model with Improvements to Rare-Word Modeling,” in Proc. of Interspeech , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
J. Li, Y. Wu, Y. Gaur et al. , “On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, “A new training pipeline for an improved neural transducer,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
Z. Dai, G. Lai, Y. Yang, and Q. Le, “Funnel-transformer: Filtering out sequential redundancy for efficient language processing,” Advances in neural information processing systems , vol. 33, pp. 4271–4282, 2020
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented transformer for speech recognition,” Proc. Interspeech 2020 , pp. 5036–5040, 2020
2020
Cited alongside, same era.
T. N. Sainath, Y. He, B. Li, A. Narayanan, R. Pang, A. Bruguier, S.-y. Chang, W. Li, R. Alvarez, Z. Chen et al. , “A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6059–6063
2020
Cited alongside, same era.
Google, “Artificial Intelligence at Google: Our Principles.” [Online]. Available: https://ai.google/principles/
Cited in the paper.
Z. Wu, D. Zhao, Q. Liang, J. Yu, A. Gulati, and R. Pang, “Dynamic sparsity neural networks for automatic speech recognition,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6014–6018
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Narayanan, T. N. Sainath, R. Pang et al. , “Cascaded encoders for unifying streaming and non-streaming ASR,” in Proc. ICASSP , 2021
2021
Later among the works it cites.
R. Botros, T. Sainath, R. David, E. Guzman, W. Li, and Y. He, “Tied & reduced rnn-t decoder,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
T. N. Sainath, Y. He, A. Narayanan et al. , “Improving the Latency and Quality of Cascaded Encoders,” in ICASSP , 2022
2022
Closest in time.