Fetching the paper…
Reading the bibliography…
We previously proposed contextual spelling correction (CSC) to correct the output of end-to-end (E2E) automatic speech recognition (ASR) models with contextual information such as name, place, etc.
“Learning small-size DNN with output-distribution-based criteria.,”
Jinyu Li, Rui Zhao, Jui-Ting Huang, and Yifan Gong, · 1914
Earlier work this paper cites.
“A light-weight contextual spelling correction model for customizing transducer-based speech recognition systems,”
Xiaoqiang Wang, Yanqing Liu, Sheng Zhao, and Jinyu Li, · 1986
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, · 2015
Earlier work this paper cites.
“Contextual speech recognition in end-to-end neural network systems using beam search.,”
Ian Williams, Anjuli Kannan, Petar S Aleksic, David Rybach, and Tara N Sainath, · 2018
Earlier work this paper cites.
“Deep context: end-to-end contextual speech recognition,”
Golan Pundak, Tara N Sainath, Rohit Prabhavalkar, Anjuli Kannan, and Ding Zhao, · 2018
Earlier work this paper cites.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Earlier work this paper cites.
“Shallow-fusion end-to-end contextual biasing,”
Ding Zhao, Tara N Sainath, David Rybach, Pat Rondon, Deepti Bhatia, Bo Li, and Ruoming Pang, · 2019
Earlier work this paper cites.
“Phoebe: Pronunciation-aware contextualization for end-to-end speech recognition,”
Antoine Bruguier, Rohit Prabhavalkar, Golan Pundak, and Tara N Sainath, · 2019
Earlier work this paper cites.
“A spelling correction model for end-to-end speech recognition,”
Jinxi Guo, Tara N Sainath, and Ron J Weiss, · 2019
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2019
Cited alongside, same era.
“Contextual RNN-T for open domain ASR,”
Mahaveer Jain, Gil Keren, Jay Mahadeokar, Geoffrey Zweig, Florian Metze, and Yatharth Saraf, · 2020
Cited alongside, same era.
“Developing RNN-T models surpassing high-performance hybrid models with customization capability,”
Jinyu Li, Rui Zhao, Zhong Meng, Yanqing Liu, Wenning Wei, Sarangarajan Parthasarathy, Vadim Mazalov, Zhenghao Wang, Lei He, Sheng Zhao, and Yifan Gong, · 2020
Cited alongside, same era.
“Improving tail performance of a deliberation e2e asr model using a large text corpus,”
Cal Peyser, Sepand Mavandadi, Tara N Sainath, James Apfel, Ruoming Pang, and Shankar Kumar, · 2020
Cited alongside, same era.
“Tree-constrained pointer generator for end-to-end contextual speech recognition,”
Guangzhi Sun, Chao Zhang, and Philip C Woodland, · 2021
Later among the works it cites.
“Delightfultts: The microsoft speech synthesis system for blizzard challenge 2021,”
Yanqing Liu, Zhihang Xu, Gang Wang, Kuan Chen, Bohan Li, Xu Tan, Jinzhu Li, Lei He, and Sheng Zhao, · 2021
Later among the works it cites.
“Transformer based deliberation for two-pass speech recognition,”
Ke Hu, Ruoming Pang, Tara N Sainath, and Trevor Strohman, · 2021
Later among the works it cites.
“Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,”
Xie Chen, Yu Wu, Zhenghao Wang, Shujie Liu, and Jinyu Li, · 2021
Later among the works it cites.
“Recent advances in end-to-end automatic speech recognition,”
Jinyu Li, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deliberation model based two-pass end-to-end speech recognition,”
Ke Hu, Tara N Sainath, Ruoming Pang, and Rohit Prabhavalkar, · 2020
Cited alongside, same era.
“Deep shallow fusion for RNN-T personalization,”
Duc Le, Gil Keren, Julian Chan, Jay Mahadeokar, Christian Fuegen, and Michael L Seltzer, · 2021
Cited alongside, same era.
“Contextualized streaming end-to-end speech recognition with trie-based deep biasing and shallow fusion,”
Duc Le, Mahaveer Jain, Gil Keren, Suyoun Kim, Yangyang Shi, Jay Mahadeokar, Julian Chan, Yuan Shangguan, Christian Fuegen, Ozlem Kalinli, Yatharth Saraf, and Michael L. Seltzer, · 2021
Cited alongside, same era.
“Instant one-shot word-learning for context-specific neural sequence-to-sequence speech recognition,”
Christian Huber, Juan Hussain, Sebastian Stüker, and Alexander Waibel, · 2021
Cited alongside, same era.
“Towards contextual spelling correction for customization of end-to-end speech recognition systems,”
Xiaoqiang Wang, Yanqing Liu, Jinyu Li, Veljko Miljanic, Sheng Zhao, and Hosam Khalil, · 2022
Later among the works it cites.
“Have best of both worlds: two-pass hybrid and E2E cascading framework for speech recognition,”
Guoli Ye, Vadim Mazalov, Jinyu Li, and Yifan Gong, · 2022
Later among the works it cites.
“Transducer-based streaming deliberation for cascaded encoders,”
Ke Hu, Tara N Sainath, Arun Narayanan, Ruoming Pang, and Trevor Strohman, · 2022
Later among the works it cites.
“Deliberation model for on-device spoken language understanding,”
Duc Le, Akshat Shrivastava, Paden Tomasello, Suyoun Kim, Aleksandr Livshits, Ozlem Kalinli, and Michael L Seltzer, · 2022
Later among the works it cites.