Fetching the paper…
Reading the bibliography…
In this paper, we construct a new Japanese speech corpus called "JTubeSpeech." Although recent end-to-end learning requires large-size speech corpora, open-sourced such corpora for languages other than English have not yet been established.
“CN-CELEB: a challenging Chinese speaker recognition dataset”, 2019
Yue Fan et al · 1911
Earlier work this paper cites.
“JNAS: Japanese speech corpus for large vocabulary continuous speech recognition research”
K. Itou et al · 1999
Earlier work this paper cites.
“Spontaneous Speech Corpus of Japanese.”
Kikuo Maekawa et al · 2000
Earlier work this paper cites.
“The Fisher corpus: A resource for the next generations of speech-to-text.”
Christopher Cieri, David Miller and Kevin Walker · 2004
Earlier work this paper cites.
“HKUST/MTS: A very large scale Mandarin telephone speech corpus”
Yi Liu et al · 2006
Earlier work this paper cites.
“Visualizing data using t-SNE”
L.v.d. Maaten and G. Hinton · 2008
Earlier work this paper cites.
“Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition”
George. Dahl et al · 2012
Earlier work this paper cites.
“Xsede: Accelerating scientific discovery computing in science & engineering, 16 (5): 62–74, sep 2014”
John Towns et al · 2014
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks”
Alex Graves and Navdeep Jaitly · 2014
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification”
E. Variani et al · 2014
Earlier work this paper cites.
“Bridges: a uniquely flexible HPC resource for new communities and data analytics”
Nicholas Nystrom et al · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books”
Vassil Panayotov et al · 2015
Cited alongside, same era.
“Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification”
Sayaka Shiota et al · 2015
Cited alongside, same era.
“YouTube-8M: A Large-Scale Video Classification Benchmark”, 2016
Sami Abu-El-Haija et al · 2016
Cited alongside, same era.
“The Kaldi OpenKWS System: Improving Low Resource Keyword Search.”
Jan Trmal et al · 2017
Cited alongside, same era.
“X-Vectors: robust DNN embeddings for speaker recognition”
D. Snyder et al · 2018
Cited alongside, same era.
“Aishell-2: Transforming mandarin asr research into industrial scale”
Shinji Watanabe et al · 2020
Later among the works it cites.
Patrick. O’Neill et al · 2021
Closest in time.
“The People’s Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage” https://openreview.net/forum?id=R8CwidgJ0yT , 2021
Daniel Galvez et al · 2021
Closest in time.
“GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio”
Guoguo Chen et al · 2021
Closest in time.
“NeMo Inverse Text Normalization: From Development To Production”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiayu Du et al · 2018
Cited alongside, same era.
“AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale”
J. Du et al · 2018
Cited alongside, same era.
“Voxceleb: Large-scale speaker verification in the wild”
Arsha Nagrani et al · 2019
Cited alongside, same era.
“Conformer: Convolution-augmented Transformer for Speech Recognition”
Anmol Gulati et al · 2020
Cited alongside, same era.
“Common Voice: A Massively-Multilingual Speech Corpus”
Rosana Ardila et al · 2020
Cited alongside, same era.
“CTC-Segmentation of Large Corpora for German End-to-End Speech Recognition”
Ludwig Kürzinger et al · 2020
Cited alongside, same era.
Yang Zhang et al · 2021
Closest in time.
“Hi-Fi Multi-Speaker English TTS Dataset”, 2021
Evelina Bakhturina et al · 2021
Closest in time.
“Recent developments on espnet toolkit boosted by conformer”
Pengcheng Guo et al · 2021
Closest in time.
“Ctc-Segmentation: Segment an audio file and obtain utterance alignments. (python package)” https://github.com/lumaku/ctc-segmentation
Ludwig Kürzinger and Dominik Winkelbauer · 2021
Closest in time.
“Num2Words: Modules to convert numbers to words.” https://github.com/savoirfairelinux/num2words
Virgil Dupras, Marius Grigaitis and Taro Ogawa · 2021
Closest in time.
“Construction of a Large-Scale Japanese ASR Corpus on TV Recordings”
Shintaro Ando and Hiromasa Fujihara · 2021
Closest in time.
“Espnet/Egs2/Laborotv/Asr1 · espnet/espnet” https://github.com/espnet/espnet/tree/master/egs2/laborotv/asr1
Various authors · 2021
Closest in time.