Fetching the paper…
Reading the bibliography…
How to leverage dynamic contextual information in end-to-end speech recognition has remained an active research area.
A. Graves, “Sequence transduction with recurrent neural networks,” in ICML Representation Learning Workshop , 2012
2012
Earlier work this paper cites.
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in Proc. ICASSP , 2015
2015
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in Proc. INTERSPEECH , 2015
2015
Earlier work this paper cites.
R. Prabhavalkar, K. Rao, T. Sainath, B. Li, L. Johnson, and N. Jaitly, “A Comparison of Sequence-to-Sequence Models for Speech Recognition,” in Proc. INTERSPEECH , 2017
2017
Earlier work this paper cites.
A. Kannan, Y. Wu, P. Nguyen, T. N. Sainath, Z. Chen, and R. Prabhavalkar, “An Analysis of Incorporating an External Language Model Into a Sequence-to-Sequence Model,” in Proc. ICASSP , 2018
2018
Earlier work this paper cites.
A. Sriram, H. Jun, S. Satheesh, and A. Coates, “Cold Fusion: Training Seq2Seq Models Together with Language Models,” in Proc. INTERSPEECH , 2018
2018
Earlier work this paper cites.
S. Toshniwal, A. Kannan, C. Chiu, Y. Wu, T. Sainath, and K. Livescu, “A Comparison of Techniques for Language Model Integration in Encoder-Decoder Speech Recognition,” in Proc. SLT , 2018
2018
Earlier work this paper cites.
G. Pundak, T. N. Sainath, R. Prabhavalkar, A. Kannan, and D. Zhao, “Deep Context: End-to-End Contextual Speech Recognition,” in Proc. ICASSP , 2018
2018
Earlier work this paper cites.
T. Kudo, “Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates,” in Proc. ACL , 2018
2018
Earlier work this paper cites.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing,” in Proc. EMNLP: System Demonstrations , 2018
2018
Cited alongside, same era.
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang, Q. Liang, D. Bhatia, Y. Shangguan, B. Li, G. Pundak, K. C. Sim, T. Bagby, S. Chang, K. Rao, and A. Gruenstein, “Streaming End-to-end Speech Recognition for Mobile Devices,” in Proc. ICASSP , 2019
2019
Cited alongside, same era.
C. Shan, C. Weng, G. Wang, D. Su, M. Luo, D. Yu, and L. Xie, “Component Fusion: Learning Replaceable Language Model Component for End-to-End Speech Recognition System,” in Proc. ICASSP , 2019
2019
Cited alongside, same era.
E. McDermott, H. Sak, and E. Variani, “A Density Ratio Approach to Language Model Fusion in End-to-End Automatic Speech Recognition,” in Proc. ASRU , 2019
2019
Cited alongside, same era.
M. Jain, G. Keren, J. Mahadeokar, G. Zweig, F. Metze, and Y. Saraf, “Contextual RNN-T For Open Domain ASR,” in Proc. INTERSPEECH , 2020
2020
Later among the works it cites.
D. Le, T. Koehler, C. Fuegen, and M. L. Seltzer, “G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR,” in Proc. ICASSP , 2020
2020
Later among the works it cites.
X. Zhang, F. Zhang, C. Liu, K. Schubert, J. Chan, P. Prakash, J. Liu, C. Yeh, F. Peng, Y. Saraf, and G. Zweig, “Benchmarking LF-MMI, CTC and RNN-T Criteria for Streaming ASR,” in Proc. SLT , 2021
2021
Closest in time.
S. Kim, Y. Shangguan, J. Mahadeokar, A. Bruguier, C. Fuegen, M. L. Seltzer, and D. Le, “Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer,” in Proc. ICASSP , 2021
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Zhao, T. N. Sainath, D. Rybach, P. Rondon, D. Bhatia, B. Li, and R. Pang, “Shallow-Fusion End-to-End Contextual Biasing,” in Proc. INTERSPEECH , 2019
2019
Cited alongside, same era.
Z. Chen, M. Jain, Y. Wang, M. L. Seltzer, and C. Fuegen, “Joint Grapheme and Phoneme Embeddings for Contextual End-to-End ASR,” in Proc. INTERSPEECH , 2019
2019
Cited alongside, same era.
D. Le, X. Zhang, W. Zheng, C. Fuegen, G. Zweig, and M. L. Seltzer, “From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition,” in Proc. ASRU , 2019
2019
Cited alongside, same era.
D. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. Cubuk, and Q. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Proc. INTERSPEECH , 2019
2019
Cited alongside, same era.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-Augmented Transformer for Speech Recognition,” in Proc. INTERSPEECH , 2020
2020
Cited alongside, same era.
E. Variani, D. Rybach, C. Allauzen, and M. Riley, “Hybrid Autoregressive Transducer (HAT),” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
2021
Closest in time.
D. Le, G. Keren, J. Chan, J. Mahadeokar, C. Fuegen, and M. L. Seltzer, “Deep Shallow Fusion for RNN-T Personalization,” in Proc. SLT , 2021
2021
Closest in time.
Y. Shi, Y. Wang, C. Wu, C. Yeh, J. Chan, F. Zhang, D. Le, and M. L. Seltzer, “Emformer: Efficient Memory Transformer Based Acoustic Model For Low Latency Streaming Speech Recognition,” in Proc. ICASSP , 2021
2021
Closest in time.
J. Mahadeokar, Y. Shangguan, D. Le, G. Keren, H. Su, T. Le, C. Yeh, C. Fuegen, and M. L. Seltzer, “Alignment Restricted Streaming Recurrent Neural Network Transducer,” in Proc. SLT , 2021
2021
Closest in time.
C. Liu, F. Zhang, D. Le, S. Kim, Y. Saraf, and G. Zweig, “Improving RNN Transducer Based ASR with Auxiliary Tasks,” in Proc. SLT , 2021
2021
Closest in time.