Fetching the paper…
Reading the bibliography…
The People's Speech is a free-to-download 30,000-hour and growing supervised conversational English speech recognition dataset licensed for academic and commercial usage under CC-BY-SA (with a CC-BY subset).
Approximate nearest neighbors: Towards removing the curse of dimensionality
P. Indyk and R. Motwani · 1998
Earlier work this paper cites.
Language recognition via ivectors and dimensionality reduction
N. Dehak, P. Torres-Carrasquillo, D. Reynolds, and R. Dehak · 2011
Earlier work this paper cites.
Langid.py: An off-the-shelf language identification tool
M. Lui and T. Baldwin · 2012
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin, 2015
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, E. Elsen, J. Engel, L. Fan, C. Fougner, T. Han, A. Hannun, B. Jun, P. LeGresley, L. Lin, S. Narang, A. Ng, S. Ozair, R. Prenger, J. Raiman, S. Satheesh, D. Seetapun, S. Sengupta, Y. Wang, Z. Wang, C. Wang, B. Xiao, D. Yogatama, J. Zhan, and Z. Zhu · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Earlier work this paper cites.
Contextual string embeddings for sequence labeling
A. Akbik, D. A. J. Blythe, and R. Vollgraf · 2018
Earlier work this paper cites.
X-vectors: Robust dnn embeddings for speaker recognition
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur · 2018
Cited alongside, same era.
pyannote.audio: neural building blocks for speaker diarization, 2019
H. Bredin, R. Yin, J. M. Coria, G. Gelly, P. Korshunov, M. Lavechin, D. Fustes, H. Titeux, W. Bouaziz, and M.-P. Gill · 2019
Cited alongside, same era.
Common voice: A massively-multilingual speech corpus, 2020
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber · 2020
Cited alongside, same era.
Gpu-accelerated viterbi exact lattice decoder for batched online and offline speech recognition, 2020
H. Braun, J. Luitjens, R. Leary, T. Kaldewey, and D. Povey · 2020
Cited alongside, same era.
Zero-shot learning in modern nlp
J. Davidson · 2020
Cited alongside, same era.
Libri-light: A benchmark for asr with limited or no supervision
Large scale weakly and semi-supervised learning for low-resource video asr, 2020
K. Singh, V. Manohar, A. Xiao, S. Edunov, R. Girshick, V. Liptchinsky, C. Fuegen, Y. Saraf, G. Zweig, and A. Mohamed · 2020
Later among the works it cites.
https://github.com/mlcommons/peoples-speech
mlcommons/peoples-speech: peoples-speech · 2021
Closest in time.
https://github.com/mozilla/DSAlign
mozilla/dsalign: Deepspeech based forced alignment tool · 2021
Closest in time.
https://docs.nvidia.com/deeplearning/nemo/user-guide/docs/en/stable/asr/models.html#conformer-ctc
Models: Conformer-ctc · 2021
Closest in time.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio
G. Chen, S. Chai, G. Wang, J. Du, W.-Q. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang, M. Jin, S. Kudanpur, S. Watanabe, S. Zhao, W. Zou, X. Li, X. Yao, Y. Wang, Z. You, and Z. Yan · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux · 2020
Cited alongside, same era.
Mls: A large-scale multilingual dataset for speech research
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert · 2020
Cited alongside, same era.
https://creativecommons.org/licenses/
Creative commons license description
Cited in the paper.
https://github.com/mozilla/DSAlign
Dsalign
Cited in the paper.
https://developer.mozilla.org/en-US/docs/Web/Guide/Audio_and_video_delivery/Adding_captions_and_subtitles_to_HTML5_video
Adding captions and subtitles to html5 video
Cited in the paper.
archive.org , a
Internet archive
Cited in the paper.
https://archive.org/services/docs/api/internetarchive/ , b
Internet archive python api
Cited in the paper.
V. J. Reddi, G. Diamos, P. Warden, P. Mattson, and D. Kanter · 2021
Closest in time.
Earnings-21: A practical benchmark for asr in the wild, 2021
M. D. Rio, N. Delworth, R. Westerman, M. Huang, N. Bhandari, J. Palakapilly, Q. McNamara, J. Dong, P. Zelasko, and M. Jette · 2021
Closest in time.