Fetching the paper…
Reading the bibliography…
As speech recognition model sizes and training data requirements grow, it is increasingly common for systems to only be available via APIs from online service providers rather than having direct access to models themselves.
2012
Earlier work this paper cites.
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, “Speaker diarization: A review of recent research,” IEEE Transactions on audio, speech, and language processing , vol. 20, no. 2, pp. 356–370, 2012
2012
Earlier work this paper cites.
H. Cucu, A. Buzo, L. Besacier, and C. Burileanu, “Statistical error correction methods for domain-specific ASR systems,” in Statistical Language and Speech Processing: First International Conference, SLSP 2013, Tarragona, Spain, July 29-31, 2013. Proceedings 1 . Springer, 2013, pp. 83–92
2013
Earlier work this paper cites.
I. Lopez-Moreno, J. Gonzalez-Dominguez, O. Plchot, D. Martinez, J. Gonzalez-Rodriguez, and P. Moreno, “Automatic language identification using deep neural networks,” in 2014 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2014, pp. 5337–5341
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
P. Bell, M. J. Gales, T. Hain, J. Kilgour, P. Lanchantin, X. Liu, A. McParland, S. Renals, O. Saz, M. Wester et al. , “The MGB challenge: Evaluating multi-genre broadcast media recognition,” in 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) . IEEE, 2015, pp. 687–693
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2016, pp. 4960–4964
2016
Earlier work this paper cites.
R. Corona, J. Thomason, and R. Mooney, “Improving black-box speech recognition using semantic parsing,” in Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers) , 2017, pp. 122–127
2017
Earlier work this paper cites.
C. Harrison, “OK Google, Siri, Alexa, Cortana; Can you tell me some stats on voice search,” Edit Agency , 2018
2018
Earlier work this paper cites.
R. Errattahi, A. El Hannani, and H. Ouahmane, “Automatic speech recognition errors detection and correction: A review,” Procedia Computer Science , vol. 128, pp. 32–37, 2018
2018
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N.-E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen et al. , “ESPnet: End-to-end speech processing toolkit,” in Proc. Interspeech 2018 , 2018, pp. 2207–2211
2018
Earlier work this paper cites.
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Esteve, “TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation,” in Speech and Computer: 20th International Conference, SPECOM 2018, Leipzig, Germany, September 18–22, 2018, Proceedings 20 . Springer, 2018, pp. 198–208
2018
Cited alongside, same era.
E. McDermott, H. Sak, and E. Variani, “A density ratio approach to language model fusion in end-to-end automatic speech recognition,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 434–441
2019
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
J. Guo, T. N. Sainath, and R. J. Weiss, “A spelling correction model for end-to-end speech recognition,” in 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5651–5655
Z. Meng, S. Parthasarathy, E. Sun, Y. Gaur, N. Kanda, L. Lu, X. Chen, R. Zhao, J. Li, and Y. Gong, “Internal language model estimation for domain-adaptive end-to-end speech recognition,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 243–250
2021
Later among the works it cites.
M. Zeineldeen, A. Glushko, W. Michel, A. Zeyer, R. Schlüter, and H. Ney, “Investigating Methods to Improve Language Model Integration for Attention-Based Encoder-Decoder ASR Models,” in Proc. Interspeech 2021 , 2021, pp. 2856–2860
2021
Later among the works it cites.
Y. Zhao, X. Yang, J. Wang, Y. Gao, C. Yan, and Y. Zhou, “BART Based Semantic Correction for Mandarin Automatic Speech Recognition System,” in Proc. Interspeech 2021 , 2021, pp. 2017–2021
2021
Later among the works it cites.
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
K. Khandelwal, P. Jyothi, A. Awasthi, and S. Sarawagi, “Black-box adaptation of ASR for accented speech,” in Proc. Interspeech 2020 , 2020, pp. 1281–1285
2020
Cited alongside, same era.
O. Hrinchuk, M. Popova, and B. Ginsburg, “Correction of automatic speech recognition with transformer sequence-to-sequence model,” in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7074–7078
2020
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Cited alongside, same era.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 7871–7880
2020
Cited alongside, same era.
J. Meyer, L. Rauchenstein, J. D. Eisenberg, and N. Howell, “Artie bias corpus: An open dataset for detecting demographic bias in speech applications,” in Proceedings of the Twelfth Language Resources and Evaluation Conference , 2020, pp. 6462–6468
2020
Cited alongside, same era.
M. Sunkara, C. Shivade, S. Bodapati, and K. Kirchhoff, “Neural inverse text normalization,” in 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 7573–7577
2021
Cited alongside, same era.
Later among the works it cites.
B. Neeley, “Speech and voice recognition market size, with 23.7% CAGR,” https://biz.crast.net/speech-and-voice-recognition-market-size-with-23-7-cagr/, November 2022
2022
Later among the works it cites.
Y. Liu, R. Ma, H. Xu, Y. He, Z. Ma, and W. Zhang, “Internal Language Model Estimation Through Explicit Context Vector Learning for Attention-based Encoder-decoder ASR,” in Proc. Interspeech 2022 , 2022, pp. 1666–1670
2022
Later among the works it cites.
K. Shen, Y. Leng, X. Tan, S. Tang, Y. Zhang, W. Liu, and E. Lin, “Mask the correct tokens: An embarrassingly simple approach for error correction,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2022, pp. 10 367–10 380
2022
Later among the works it cites.
Z. Gekhman, D. Zverinski, J. Mallinson, and G. Beryozkin, “RED-ACE: Robust error detection for ASR using confidence embeddings,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2022, pp. 2800–2808
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Closest in time.