Fetching the paper…
Reading the bibliography…
Text-only and semi-supervised training based on audio-only data has gained popularity recently due to the wide availability of unlabeled text and speech data.
M. Schuster and K. Nakajima, “Japanese and Korean voice search,” in Proc. IEEE ICASSP , 2012, pp. 5149–5152
2012
Earlier work this paper cites.
H. Liao, E. McDermott, and A. Senior, “Large scale deep neural network acoustic modeling with semi-supervised training data for YouTube video transcription,” in ASRU , 2013, pp. 368–373
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
X. Gonzalvo, S. Tazari, C.-a. Chan, M. Becker, A. Gutkin, and H. Silen, “Recent advances in google real-time HMM-driven unit selection synthesizer,” Interspeech , 2016
2016
Earlier work this paper cites.
G. Pundak and T. Sainath, “Lower frame rate neural network acoustic models,” in Proc. Interspeech 2016 , 2016, pp. 22–26
2016
Earlier work this paper cites.
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. N. Sainath, and M. Bacchiani, “Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google home,” in Proc. Interspeech , 2017, pp. 379–383
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Shin, Y. Lee, and K. Jung, “Effective sentence scoring method using BERT for speech recognition,” in ACML . PMLR, 2019, pp. 1081–1093
2019
Earlier work this paper cites.
J. Guo, T. N. Sainath, and R. J. Weiss, “A spelling correction model for end-to-end speech recognition,” in Proc. IEEE ICASSP , 2019, pp. 5651–5655
2019
Earlier work this paper cites.
P. Bahar, T. Bieschke, and H. Ney, “A comparative study on end-to-end speech to text translation,” in ASRU . IEEE, 2019, pp. 792–799
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Narayanan, R. Prabhavalkar, C.-C. Chiu, D. Rybach, T. N. Sainath, and T. Strohman, “Recognizing long-form speech using streaming end-to-end models,” in ASRU , 2019, pp. 920–927
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
T. N. Sainath, R. Pang, D. Rybach, Y. He, R. Prabhavalkar, W. Li, M. Visontai, Q. Liang, T. Strohman, Y. Wu, I. McGraw, and C.-C. Chiu, “Two-pass end-to-end speech recognition,” in Proc. Interspeech , 2019, pp. 2773–2777
2019
Cited alongside, same era.
2020
Cited alongside, same era.
K. Hu, T. N. Sainath, R. Pang, and R. Prabhavalkar, “Deliberation model based two-pass end-to-end speech recognition,” in IEEE ICASSP , 2020, pp. 7799–7803
2020
Cited alongside, same era.
2021
Later among the works it cites.
K. Hu, R. Pang, T. N. Sainath, and T. Strohman, “Transformer based deliberation for two-pass speech recognition,” in SLT , 2021, pp. 68–74
2021
Later among the works it cites.
S. Mavandadi, T. N. Sainath, K. Hu, and Z. Wu, “A deliberation-based joint acoustic and text decoder,” in Interspeech , 2021
2021
Later among the works it cites.
D. Fohr and I. Illina, “BERT-based semantic model for rescoring n-best speech recognition list,” in INTERSPEECH , 2021
2021
Later among the works it cites.
Y. Tang, J. Pino, C. Wang, X. Ma, and D. Genzel, “A general multi-task learning framework to leverage text data for speech to text tasks,” in IEEE ICASSP , 2021, pp. 6209–6213
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. Hrinchuk, M. Popova, and B. Ginsburg, “Correction of automatic speech recognition with transformer sequence-to-sequence model,” in IEEE ICASSP , 2020, pp. 7074–7078
2020
Cited alongside, same era.
T. N. Sainath, R. Pang, R. J. Weiss, Y. He, C.-c. Chiu, and T. Strohman, “An attention-based joint acoustic and text on-device end-to-end model,” in IEEE ICASSP , 2020, pp. 7039–7043
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T. N. Sainath, Y. He, A. Narayanan, R. Botros, R. Pang, D. Rybach, C. Allauzen, E. Variani, J. Qin, Q.-N. Le-The, S.-Y. Chang, B. Li, A. Gulati, C.-C. Yu, Jiahui Chiu, D. Caseiro, W. Li, Q. Liang, P. Rondo et al. , “An efficient streaming non-recurrent on-device end-to-end model with improvements to rare-word modeling,” Interspeech , 2021
2021
Cited alongside, same era.
A. Narayanan, T. N. Sainath, R. Pang, J. Yu, C.-C. Chiu, R. Prabhavalkar, E. Variani, and T. Strohman, “Cascaded encoders for unifying streaming and non-streaming asr,” in IEEE ICASSP , 2021, pp. 5629–5633
2021
Cited alongside, same era.
X. Chen, Y. Wu, Z. Wang, S. Liu, and J. Li, “Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,” in IEEE ICASSP , 2021, pp. 5904–5908
2021
Cited alongside, same era.
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
T. Doutre, W. Han, M. Ma, Z. Lu, C.-C. Chiu, R. Pang, A. Narayanan, A. Misra, Y. Zhang, and L. Cao, “Improving streaming automatic speech recognition with non-streaming model distillation on unsupervised data,” in IEEE ICASSP 2021 , 2021, pp. 6558–6562
2021
Later among the works it cites.
2021
Later among the works it cites.
B. Li, A. Gulati, J. Yu, T. N. Sainath, C.-C. Chiu, A. Narayanan, S.-Y. Chang, R. Pang, Y. He, J. Qin et al. , “A better and faster end-to-end model for streaming ASR,” in IEEE ICASSP , 2021, pp. 5634–5638
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Hu, T. N. Sainath, R. Pang, and R. Prabhavalkar, “Transducer-based streaming deliberation for cascaded encoders,” in IEEE ICASSP 2022 (to appear)
2022
Closest in time.
T. N. Sainath, Y. R. He, A. Narayanan, R. Botros, W. Wang, D. Qiu, C.-C. Chiu, R. Prabhavalkar, A. Gruenstein, A. Gulati, B. Li, D. Rybach, E. Guzman, I. McGraw, J. Qin, K. Choromanski, Q. Liang, R. David, R. Pang, S.-y. Chang, T. Strohman, W. R. Huang, W. Han, Y. Wu, and Y. Zhang, “Improving the latency and quality of cascaded encoders,” IEEE ICASSP , 2022 (to appear)
2022
Closest in time.