Fetching the paper…
Reading the bibliography…
Commonly used speech corpora inadequately challenge academic and commercial ASR systems.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in Acoustics, Speech, and Signal Processing, IEEE International Conference on , vol. 1. IEEE Computer Society, 1992, pp. 517–520
1992
Earlier work this paper cites.
E. E. Shriberg, “Preliminaries to a theory of speech disfluencies,” 1994
1994
Earlier work this paper cites.
P. A. Heeman and J. Allen, “Speech repains, intonational phrases, and discourse markers: modeling speakers’ utterances in spoken dialogue,” 1999
1999
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The fisher corpus: a resource for the next generations of speech-to-text.” in LREC , vol. 4, 2004, pp. 69–71
2004
Earlier work this paper cites.
J. Carletta, “Announcing the ami meeting corpus,” The ELRA Newsletter , vol. 11, no. 1, pp. 3–5, 2006
2006
Earlier work this paper cites.
J. Fiscus, J. Ajot, and J. Garofolo, “The rich transcription 2007 meeting recognition evaluation.” The Joint Proceedings of the 2006 CLEAR and RT Evaluations, 2007
2007
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding , no. CONF. IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
2012
Cited alongside, same era.
2015
Cited alongside, same era.
X. Cui, V. Goel, and B. Kingsbury, “Data augmentation for deep neural network acoustic modeling,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 9, pp. 1469–1477, 2015
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
K. J. Han, A. Chandrashekaran, J. Kim, and I. Lane, “Densely connected networks for conversational speech recognition.” in INTERSPEECH , 2018, pp. 796–800
2018
Later among the works it cites.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-end speech processing toolkit,” in Proceedings of Interspeech , 2018, pp. 2207–2211. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2018-1456
2018
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
V. Peddinti, Y. Wang, D. Povey, and S. Khudanpur, “Low latency acoustic modeling using temporal convolution and lstms,” vol. 25, no. 3. IEEE, 2017, pp. 373–377
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid ctc/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing , vol. 11, no. 8, pp. 1240–1253, 2017
2017
Cited alongside, same era.
2020
Later among the works it cites.
P. Szymański, P. Żelasko, M. Morzy, A. Szymczak, M. Żyła-Hoppe, J. Banaszczak, L. Augustyniak, J. Mizgajski, and Y. Carmiel, “WER we are and WER we think we are,” in Findings of the Association for Computational Linguistics: EMNLP 2020 . Online: Association for Computational Linguistics, Nov. 2020, pp. 3290–3295. [Online]. Available: https://www.aclweb.org/anthology/2020.findings-emnlp.295
2020
Later among the works it cites.