Fetching the paper…
Reading the bibliography…
End-to-end automatic speech recognition (ASR) systems, such as recurrent neural network transducer (RNN-T), have become popular, but rare word remains a challenge.
“Improved backing-off for m-gram language modeling,”
Reinhard Kneser and Hermann Ney, · 1995
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Nagendra Goel, Mirko Hannemann, Yanmin Qian, Petr Schwarz, and Georg Stemmer, · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“On using monolingual corpora in neural machine translation,”
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Composition-based on-the-fly rescoring for salient n-gram biasing,”
Keith Hall, Eunjoon Cho, Cyril Allauzen, Francoise Beaufays, Noah Coccaro, Kaisuke Nakajima, Michael Riley, Brian Roark, David Rybach, and Linda Zhang, · 2015
Earlier work this paper cites.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2016
Earlier work this paper cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Earlier work this paper cites.
“Cold fusion: Training seq2seq models together with language models,”
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates, · 2017
Earlier work this paper cites.
“Acoustic modeling for google home.,”
Bo Li, Tara N Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K Chin, et al., · 2017
Cited alongside, same era.
“Contextual speech recognition in end-to-end neural network systems using beam search.,”
Ian Williams, Anjuli Kannan, Petar S Aleksic, David Rybach, and Tara N Sainath, · 2018
Cited alongside, same era.
Taku Kudo, · 2018
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Cited alongside, same era.
“Semi-supervised training for end-to-end models via weak distillation,”
Bo Li, Tara N Sainath, Ruoming Pang, and Zelin Wu, · 2019
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Later among the works it cites.
“Improving tail performance of a deliberation e2e asr model using a large text corpus,”
Cal Peyser, Sepand Mavandadi, Tara N Sainath, James Apfel, Ruoming Pang, and Shankar Kumar, · 2020
Closest in time.
“Developing rnn-t models surpassing high-performance hybrid models with customization capability,”
Jinyu Li, Rui Zhao, Zhong Meng, Yanqing Liu, Wenning Wei, Sarangarajan Parthasarathy, Vadim Mazalov, Zhenghao Wang, Lei He, Sheng Zhao, et al., · 2020
Closest in time.
“Efficient minimum word error rate training of rnn-transducer for end-to-end speech recognition,”
Jinxi Guo, Gautam Tiwari, Jasha Droppo, Maarten Van Segbroeck, Che-Wei Huang, Andreas Stolcke, and Roland Maas, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tara N Sainath, Ruoming Pang, David Rybach, Yanzhang He, Rohit Prabhavalkar, Wei Li, Mirkó Visontai, Qiao Liang, Trevor Strohman, Yonghui Wu, et al., · 2019
Cited alongside, same era.
“Shallow-fusion end-to-end contextual biasing.,”
Ding Zhao, Tara N Sainath, David Rybach, Pat Rondon, Deepti Bhatia, Bo Li, and Ruoming Pang, · 2019
Cited alongside, same era.
“Scalable multi corpora neural language models for asr,”
Anirudh Raju, Denis Filimonov, Gautam Tiwari, Guitang Lan, and Ariya Rastrow, · 2019
Cited alongside, same era.
“Towards fast and accurate streaming end-to-end asr,”
Bo Li, Shuo-yiin Chang, Tara N Sainath, Ruoming Pang, Yanzhang He, Trevor Strohman, and Yonghui Wu, · 2020
Closest in time.
“Deliberation model based two-pass end-to-end speech recognition,”
Ke Hu, Tara N Sainath, Ruoming Pang, and Rohit Prabhavalkar, · 2020
Closest in time.
“Improving proper noun recognition in end-to-end asr by customization of the mwer loss criterion,”
Cal Peyser, Tara N Sainath, and Golan Pundak, · 2020
Closest in time.
“Class lm and word mapping for contextual biasing in end-to-end asr,”
Rongqing Huang, Ossama Abdel-hamid, Xinwei Li, and Gunnar Evermann, · 2020
Closest in time.