Fetching the paper…
Reading the bibliography…
Contextual biasing is an important and challenging task for end-to-end automatic speech recognition (ASR) systems, which aims to achieve better recognition performance by biasing the ASR system to particular context phrases such as person names, music list, proper nouns, etc.
J. Li, R. Zhao, J.-T. Huang, and Y. Gong, “Learning small-size DNN with output-distribution-based criteria.” in Proc. Interspeech , 2014, pp. 1910–1914
1914
Earlier work this paper cites.
X. Wang, Y. Liu, S. Zhao, and J. Li, “A light-weight contextual spelling correction model for customizing transducer-based speech recognition systems,” in Proc. Interspeech , 2021, pp. 1982–1986
1986
Earlier work this paper cites.
D. Nadeau and S. Sekine, “A survey of named entity recognition and classification,” Lingvisticae Investigationes: 30.1 , pp. 3–26, 2007
2007
Earlier work this paper cites.
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in Proc. EMNLP , 2014, pp. 1724–1734
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Proc. NeurIPS , vol. 27, 2014, pp. 3104–3112
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proc. ICLR , 2015
2015
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in NeurIPS Deep Learning and Representation Learning Workshop , 2015
2015
Earlier work this paper cites.
T. Mitchell, W. Cohen, E. Hruschka, P. Talukdar, B. Yang, J. Betteridge, A. Carlson, B. Dalvi, M. Gardner, B. Kisiel et al. , “Never-ending learning,” in Proc. AAAI , 2015, pp. 2302–2310
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Proc. ICASSP . IEEE, 2016, pp. 4960–4964
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” in NeurIPS Deep Learning Symposium , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NeurIPS , 2017, pp. 6000–6010
2017
Earlier work this paper cites.
I. Williams, A. Kannan, P. S. Aleksic, D. Rybach, and T. N. Sainath, “Contextual speech recognition in end-to-end neural network systems using beam search.” in Proc. Interspeech , 2018, pp. 2227–2231
2018
Earlier work this paper cites.
G. Pundak, T. N. Sainath, R. Prabhavalkar, A. Kannan, and D. Zhao, “Deep context: end-to-end contextual speech recognition,” in Proc. SLT . IEEE, 2018, pp. 418–425
2018
Earlier work this paper cites.
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, and R. Pang, “Streaming end-to-end speech recognition for mobile devices,” in Proc. ICASSP . IEEE, 2019, pp. 6381–6385
2019
Earlier work this paper cites.
J. Li, R. Zhao, H. Hu, and Y. Gong, “Improving rnn transducer modeling for end-to-end speech recognition,” in Proc. ASRU . IEEE, 2019, pp. 114–121
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang et al. , “Streaming end-to-end speech recognition for mobile devices,” in Proc. ICASSP . IEEE, 2019, pp. 6381–6385
2019
Cited alongside, same era.
D. Zhao, T. N. Sainath, D. Rybach, D. Bhatia, B. Li, and R. Pang, “Shallow-fusion end-to-end contextual biasing,” in Proc. Interspeech , 2019, pp. 1418–1422
2019
Cited alongside, same era.
O. Hrinchuk, M. Popova, and B. Ginsburg, “Correction of automatic speech recognition with transformer sequence-to-sequence model,” in Proc. ICASSP . IEEE, 2020, pp. 7074–7078
2020
Later among the works it cites.
K. Hu, T. N. Sainath, R. Pang, and R. Prabhavalkar, “Deliberation model based two-pass end-to-end speech recognition,” in Proc. ICASSP . IEEE, 2020, pp. 7799–7803
2020
Later among the works it cites.
J. Li, A. Sun, J. Han, and C. Li, “A survey on deep learning for named entity recognition,” IEEE Transactions on Knowledge and Data Engineering , vol. 34, no. 1, pp. 50–70, 2020
2020
Later among the works it cites.
T. Pellissier Tanon, G. Weikum, and F. Suchanek, “Yago 4: A reason-able knowledge base,” in Proc. ESWC , 2020, pp. 583–596
2020
Later among the works it cites.
X. Chen, Y. Wu, Z. Wang, S. Liu, and J. Li, “Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,” in Proc. ICASSP . IEEE, 2021, pp. 5904–5908
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Bruguier, R. Prabhavalkar, G. Pundak, and T. N. Sainath, “Phoebe: Pronunciation-aware contextualization for end-to-end speech recognition,” in Proc. ICASSP . IEEE, 2019, pp. 6171–6175
2019
Cited alongside, same era.
J. Guo, T. N. Sainath, and R. J. Weiss, “A spelling correction model for end-to-end speech recognition,” in Proc. ICASSP , 2019, pp. 5651–5655
2019
Cited alongside, same era.
S. Zhang, M. Lei, and Z. Yan, “Investigation of transformer based spelling correction model for CTC-based end-to-end Mandarin speech recognition,” in Proc. Interspeech , 2019, pp. 2180–2184
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
T. N. Sainath, Y. He, B. Li, A. Narayanan, R. Pang, A. Bruguier, S.-y. Chang, W. Li, R. Alvarez, Z. Chen et al. , “A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,” in Proc. ICASSP . IEEE, 2020, pp. 6059–6063
2020
Cited alongside, same era.
J. Li, R. Zhao, Z. Meng, Y. Liu, W. Wei, S. Parthasarathy, V. Mazalov, Z. Wang, L. He, S. Zhao, and Y. Gong, “Developing RNN-T models surpassing high-performance hybrid models with customization capability,” in Proc. Interspeech , 2020, pp. 3590–3594
2020
Cited alongside, same era.
G. Saon, Z. Tüske, and K. Audhkhasi, “Alignment-length synchronous decoding for RNN transducer,” in Proc. ICASSP . IEEE, 2020, pp. 7804–7808
2020
Cited alongside, same era.
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, “A new training pipeline for an improved neural transducer,” in Proc. Interspeech , 2020, pp. 2812–2816
2020
Cited alongside, same era.
2021
Later among the works it cites.
D. Le, G. Keren, J. Chan, J. Mahadeokar, C. Fuegen, and M. L. Seltzer, “Deep shallow fusion for RNN-T personalization,” in Proc. SLT . IEEE, 2021, pp. 251–257
2021
Later among the works it cites.
D. Le, M. Jain, G. Keren, S. Kim, Y. Shi, J. Mahadeokar, J. Chan, Y. Shangguan, C. Fuegen, O. Kalinli, Y. Saraf, and M. L. Seltzer, “Contextualized streaming end-to-end speech recognition with trie-based deep biasing and shallow fusion,” in Proc. Interspeech , 2021, pp. 1772–1776
2021
Later among the works it cites.
C. Huber, J. Hussain, S. Stüker, and A. Waibel, “Instant one-shot word-learning for context-specific neural sequence-to-sequence speech recognition,” in Proc. ASRU . IEEE, 2021, pp. 1–7
2021
Later among the works it cites.
G. Sun, C. Zhang, and P. C. Woodland, “Tree-constrained pointer generator for end-to-end contextual speech recognition,” in Proc. ASRU . IEEE, 2021, pp. 780–787
2021
Later among the works it cites.
Y. Leng, X. Tan, L. Zhu, J. Xu, R. Luo, L. Liu, T. Qin, X. Li, E. Lin, and T.-Y. Liu, “Fastcorrect: Fast error correction with edit alignment for automatic speech recognition,” in Proc. NeurIPS , vol. 34, 2021, pp. 21 708–21 719
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Li, “Recent advances in end-to-end automatic speech recognition,” APSIPA Transactions on Signal and Information Processing , vol. 11, no. 1, 2022
2022
Closest in time.
G. Ye, V. Mazalov, J. Li, and Y. Gong, “Have best of both worlds: two-pass hybrid and e2e cascading framework for speech recognition,” in Proc. ICASSP . IEEE, 2022, pp. 7432–7436
2022
Closest in time.
W. Wang, K. Hu, and T. N. Sainath, “Deliberation of streaming rnn-transducer by non-autoregressive decoding,” in Proc. ICASSP . IEEE, 2022, pp. 7452–7456
2022
Closest in time.