Fetching the paper…
Reading the bibliography…
Neural transducer is now the most popular end-to-end model for speech recognition, due to its naturally streaming ability.
“Long Short-Term Memory,”
Sepp Hochreiter and Jurgen Schmidhuber, · 1997
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jurgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Improving wideband speech recognition using mixed-bandwidth training data in CD-DNN-HMM,”
Jinyu Li, Dong Yu, Jui-Ting Huang, and Yifan Gong, · 2012
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Exploring neural transducers for end-to-end speech recognition,”
Eric Battenberg, Jitong Chen, Rewon Child, Adam Coates, Yashesh Gaur, Yi Li, Hairong Liu, and et al., · 2017
Earlier work this paper cites.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi, · 2017
Earlier work this paper cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, et al., · 2018
Earlier work this paper cites.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N. Sainath, Zhifeng Chen, and Rohit Prabhavalkar, · 2018
Earlier work this paper cites.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, et al., · 2019
Earlier work this paper cites.
“Transformer-transducer: End-to-end speech recognition with self-attention,”
Ching-Feng Yeh, Jay Mahadeokar, et al., · 2019
Earlier work this paper cites.
“Improving RNN transducer modeling for end-to-end speech recognition,”
Jinyu Li, Rui Zhao, Hu Hu, and Yifan Gong, · 2019
Earlier work this paper cites.
“Personalization of end-to-end speech recognition on mobile devices for named entities,”
Khe Chai Sim, Francoise Beaufays, et al., · 2019
Cited alongside, same era.
“A density ratio approach to language model fusion in end-to-end automatic speech recognition,”
Erik McDermott, Hasim Sak, and Ehsan Variani, · 2019
Cited alongside, same era.
“Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,”
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar, · 2020
Cited alongside, same era.
“On the comparison of popular end-to-end models for large scale speech recognition,”
Jinyu Li, Yu Wu, Yashesh Gaur, Chengyi Wang, Rui Zhao, and Shujie Liu, · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, and et al., · 2020
Cited alongside, same era.
“Improving RNN-T for domain scaling using semi-supervised training with neural TTS,”
Yan Deng, Rui Zhao, Zhong Meng, Xie Chen, Bin Liu, Jinyu Li, Yifan Gong, and Lei He, · 2021
Later among the works it cites.
“Using synthetic audio to improve the recognition of out-of-vocabulary words in end-to-end ASR systems,”
Xianrui Zheng, Yulan Liu, Deniz Gunceler, and Daniel Willett, · 2021
Later among the works it cites.
“SynthASR: Unlocking Synthetic Data for Speech Recognition,”
Amin Fazel, Wei Yang, Yulan Liu, Roberto Barra-Chicote, Yixiong Meng, Roland Maas, and Jasha Droppo, · 2021
Later among the works it cites.
“On addressing practical challenges for RNN-Transducer,”
Rui Zhao, Jian Xue, Jinyu Li, Wenning Wei, Lei He, and Yifan Gong, · 2021
Later among the works it cites.
“Language model fusion for streaming end to end speech recognition,”
Rodrigo Cabrera, Xiaofeng Liu, Mohammadreza Ghodsi, Zebulun Matteson, Eugene Weinstein, and Anjuli Kannan, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,”
Tara N. Sainath, Yanzhang He, Bo Li, Arun Narayanan, Ruoming Pang, Antoine Bruguier, Shuo yiin Chang, and et al., · 2020
Cited alongside, same era.
“Developing RNN-T models surpassing high-performance hybrid models with customization capability,”
Jinyu Li, Rui Zhao, Zhong Meng, Yanqing Liu, Wenning Wei, Sarangarajan Parthasarathy, Vadim Mazalov, and et. al., · 2020
Cited alongside, same era.
“Improving accuracy of rare words for rnn-transducer through unigram shallow fusion,”
Vijay Ravi, Yile Gu, Ankur Gandhe, Ariya Rastrow, Linda Liu, Denis Filimonov, Scott Novotney, and Ivan Bulyko, · 2020
Cited alongside, same era.
“Contextualized streaming end-to-end speech recognition with trie-based deep biasing and shallow fusion,”
Duc Le, Mahaveer Jain, Suyoun Kim Gil Keren, Yangyang Shi, Julian Chan Jay Mahadeokar, and et. al., · 2020
Cited alongside, same era.
“Hybrid autoregressive transducer (HAT),”
Ehsan Variani, David Rybach, Cyril Allauzen, and Michael Riley, · 2020
Cited alongside, same era.
“Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,”
Xie Chen, Yu Wu, Zhenghao Wang, Shujie Liu, and Jinyu Li, · 2021
Cited alongside, same era.
“Internal language model estimation for domain-adaptive end-to-end speech recognition,”
Zhong Meng, Sarangarajan Parthasarathy, Eric Sun, Yashesh Gaur, Naoyuki Kanda, Liang Lu, Xie Chen, and et. al., · 2021
Later among the works it cites.
“Internal language model training for domain-adaptive end-to-end speech recognition,”
Zhong Meng, Naoyuki Kanda, Yashesh Gaur, Sarangarajan Parthasarathy, Eric Sun, Liang Lu, Xie Chen, and et. al., · 2021
Later among the works it cites.
“Recent advances in end-to-end automatic speech recognition,”
Jinyu Li, · 2022
Closest in time.
“A likelihood ratio-based domain adaptation method for end-to-end models,”
Chhavi Choudhury, Ankur Gandhe, Xiaohan Ding, and Ivan Bulyko, · 2022
Closest in time.
“Adaptive discounting of implicit language models in rnn-transducers,”
Vinit Unni, Shreya Khare, Ashish Mittal, Preethi Jyothi, Sunita Sarawagi, and Samarth Bharadwaj, · 2022
Closest in time.
“Factorized neural transducer for efficient language model adaptation,”
Xie Chen, Zhong Meng, Sarangarajan Parthasarathy, and Jinyu Li, · 2022
Closest in time.