Fetching the paper…
Reading the bibliography…
Attention-based contextual biasing approaches have shown significant improvements in the recognition of generic and/or personal rare-words in End-to-End Automatic Speech Recognition (E2E ASR) systems like neural transducers.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Learning personalized pronunciations for contact names recognition,”
Tony Bruguier, Fuchun Peng, and Françoise Beaufays, · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Pointer sentinel mixture models,”
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher, · 2017
Earlier work this paper cites.
“Deep context: end-to-end contextual speech recognition,”
Golan Pundak, Tara N Sainath, Rohit Prabhavalkar, Anjuli Kannan, and Ding Zhao, · 2018
Earlier work this paper cites.
“No need for a lexicon? evaluating the value of the pronunciation lexica in end-to-end models,”
Tara N Sainath, Rohit Prabhavalkar, Shankar Kumar, Seungji Lee, Anjuli Kannan, David Rybach, Vlad Schogol, Patrick Nguyen, Bo Li, Yonghui Wu, et al., · 2018
Earlier work this paper cites.
“Subword regularization: Improving neural network translation models with multiple subword candidates,”
Taku Kudo, · 2018
Earlier work this paper cites.
“Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
“Shallow-fusion end-to-end contextual biasing.,”
Ding Zhao, Tara N Sainath, David Rybach, Pat Rondon, Deepti Bhatia, Bo Li, and Ruoming Pang, · 2019
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Cited alongside, same era.
“Shallow-Fusion End-to-End Contextual Biasing,”
Ding Zhao, Tara N. Sainath, David Rybach, Pat Rondon, Deepti Bhatia, Bo Li, and Ruoming Pang, · 2019
Cited alongside, same era.
“Joint grapheme and phoneme embeddings for contextual end-to-end asr,”
Zhehuai Chen, Mahaveer Jain, Yongqiang Wang, Michael L Seltzer, and Christian Fuegen, · 2019
Cited alongside, same era.
“Deep shallow fusion for rnn-t personalization,”
Duc Le, Gil Keren, Julian Chan, Jay Mahadeokar, Christian Fuegen, and Michael L Seltzer, · 2021
Later among the works it cites.
“Personalization strategies for end-to-end speech recognition systems,”
Aditya Gourav, Linda Liu, Ankur Gandhe, Yile Gu, Guitang Lan, Xiangyang Huang, Shashank Kalmane, Gautam Tiwari, Denis Filimonov, Ariya Rastrow, et al., · 2021
Later among the works it cites.
“Context-aware transformer transducer for speech recognition,”
Feng-Ju Chang, Jing Liu, Martin Radfar, Athanasios Mouchtaris, Maurizio Omologo, Ariya Rastrow, and Siegfried Kunzmann, · 2021
Later among the works it cites.
“Adapting long context nlm for asr rescoring in conversational agents,”
Ashish Shenoy, Sravan Bodapati, Monica Sunkara, Srikanth Ronanki, and Katrin Kirchhoff, · 2021
Later among the works it cites.
Duc Le, Mahaveer Jain, Gil Keren, Suyoun Kim, Yangyang Shi, Jay Mahadeokar, Julian Chan, Yuan Shangguan, Christian Fuegen, Ozlem Kalinli, et al., · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Phoebe: Pronunciation-aware contextualization for end-to-end speech recognition,”
Antoine Bruguier, Rohit Prabhavalkar, Golan Pundak, and Tara N Sainath, · 2019
Cited alongside, same era.
“Contextual rnn-t for open domain asr,”
Mahaveer Jain, Gil Keren, Jay Mahadeokar, Geoffrey Zweig, Florian Metze, and Yatharth Saraf, · 2020
Cited alongside, same era.
“Rnn-transducer with stateless prediction network,”
Mohammadreza Ghodsi, Xiaofeng Liu, James Apfel, Rodrigo Cabrera, and Eugene Weinstein, · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2020
Cited alongside, same era.
Later among the works it cites.
“Contextual adapters for personalized speech recognition in neural transducers,”
Kanthashree Mysore Sathyendra, Thejaswi Muniyappa, Feng-Ju Chang, Jing Liu, Jinru Su, Grant P Strimel, Athanasios Mouchtaris, and Siegfried Kunzmann, · 2022
Later among the works it cites.
“Improving end-to-end contextual speech recognition with fine-grained contextual knowledge selection,”
Minglun Han, Linhao Dong, Zhenlin Liang, Meng Cai, Shiyu Zhou, Zejun Ma, and Bo Xu, · 2022
Later among the works it cites.
“Fast contextual adaptation with neural associative memory for on-device personalized speech recognition,”
Tsendsuren Munkhdalai, Khe Chai Sim, Angad Chandorkar, Fan Gao, Mason Chua, Trevor Strohman, and Françoise Beaufays, · 2022
Later among the works it cites.