Fetching the paper…
Reading the bibliography…
Measuring automatic speech recognition (ASR) system quality is critical for creating user-satisfying voice-driven applications.
M. J. Hunt, “Figures of merit for assessing connected-word recognisers,” Speech Communication , vol. 9, no. 4, pp. 329–336, 1990
1990
Earlier work this paper cites.
J. S. Garofolo, E. M. Voorhees, C. G. Auzanne, V. M. Stanford, and B. A. Lund, “1998 TREC-7 spoken document retrieval track overview and results,” NIST SPECIAL PUBLICATION SP , pp. 79–90, 1999
1999
Earlier work this paper cites.
J. Makhoul, F. Kubala, R. Schwartz, R. Weischedel et al. , “Performance measures for information extraction,” in Proceedings of DARPA broadcast news workshop . Herndon, VA, 1999, pp. 249–252
1999
Earlier work this paper cites.
I. A. McCowan, D. Moore, J. Dines, D. Gatica-Perez, M. Flynn, P. Wellner, and H. Bourlard, “On the use of information retrieval measures for speech recognition evaluation,” IDIAP, Tech. Rep., 2004
2004
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” in ICML Representation Learning Workshop , 2012
2012
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in NAACL , 2018
2018
Earlier work this paper cites.
S. Gupta, R. Shah, M. Mohit, A. Kumar, and M. Lewis, “Semantic parsing for task oriented dialog using hierarchical representations,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, Eds. Association for Computational Linguistics, 2018, pp. 2787–2792. [Online]. Available: https://doi.org/10.18653/v1/d18-1300
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Later among the works it cites.
A. Aghajanyan, J. Maillard, A. Shrivastava, K. Diedrick, M. Haeger, H. Li, Y. Mehdad, V. Stoyanov, A. Kumar, M. Lewis, and S. Gupta, “Conversational semantic parsing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 , B. Webber, T. Cohn, Y. He, and Y. Liu, Eds. Association for Computational Linguistics, 2020, pp. 5026–5035. [Online]. Available: https://doi.org/10.18653/v1/2020.emnlp-main.408
2020
Later among the works it cites.
W. Yuan, G. Neubig, and P. Liu, “BARTscore: Evaluating generated text as text generation,” Advances in Neural Information Processing Systems , vol. 34, pp. 27 263–27 277, 2021
2021
Closest in time.
S. Kim, A. Arora, D. Le, C.-F. Yeh, C. Fuegen, O. Kalinli, and M. L. Seltzer, “Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding,” in Proc. INTERSPEECH , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Closest in time.
Y. Shi, Y. Wang, C. Wu, C. Yeh, J. Chan, F. Zhang, D. Le, and M. L. Seltzer, “Emformer: Efficient Memory Transformer Based Acoustic Model For Low Latency Streaming Speech Recognition,” in Proc. ICASSP , 2021
2021
Closest in time.
J. Mahadeokar, Y. Shangguan, D. Le, G. Keren, H. Su, T. Le, C. Yeh, C. Fuegen, and M. L. Seltzer, “Alignment Restricted Streaming Recurrent Neural Network Transducer,” in Proc. SLT , 2021
2021
Closest in time.
D. Le, M. Jain, G. Keren, S. Kim, Y. Shi, J. Mahadeokar, J. Chan, Y. Shangguan, C. Fuegen, O. Kalinli, Y. Saraf, and M. L. Seltzer, “Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion,” in Proc. Interspeech , 2021, pp. 1772–1776
2021
Closest in time.