Fetching the paper…
Reading the bibliography…
This paper introduces a new multi-speaker English dataset for training text-to-speech models.
D. B. Paul and J. Baker, “The design for the Wall Street Journal-based CSR corpus,” in Speech and Natural Language: Proceedings of a Workshop Held at Harriman, New York, February 23-26, 1992 , 1992
1992
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” in ICASSP , 1992
1992
Earlier work this paper cites.
J. L. Hintze and R. D. Nelson, “Violin plots: a box plot-density trace synergism,” The American Statistician , vol. 52, no. 2, pp. 181–184, 1998
1998
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The Fisher corpus: a resource for the next generations of speech-to-text.” in LREC , 2004
2004
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in ICML , 2006
2006
Earlier work this paper cites.
P. Taylor, “Text-to-Speech Synthesis,” in Cambridge university press , 2009
2009
Earlier work this paper cites.
S. King and V. Karaiskos, “The blizzard challenge 2013,” 2013
2013
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in ICASSP , 2015
2015
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
Munich Artificial Intelligence Laboratories GmbH, “The m-ailabs speech dataset,” https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/ , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Cited alongside, same era.
J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92),” 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
R. Valle, J. Li, R. Prenger, and B. Catanzaro, “Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,” in ICASSP , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. X. Koh, A. Mislan, K. Khoo, B. Ang, W. Ang, C. Ng, and Y. Tan, “Building the singapore english national speech corpus,” Malay , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
“LibriVox - Free public domain audiobooks,” https://librivox.org/
Cited in the paper.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux, “Libri-Light: A benchmark for ASR with limited or no supervision,” in ICASSP , 2020
2020
Later among the works it cites.
S. Kriman, S. Beliaev, B. Ginsburg, J. Huang, O. Kuchaiev, V. Lavrukhin, R. Leary, J. Li, and Y. Zhang, “Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions,” in ICASSP , 2020
2020
Later among the works it cites.
L. Kürzinger, D. Winkelbauer, L. Li, T. Watzel, and G. Rigoll, “CTC-Segmentation of Large Corpora for German End-to-End Speech Recognition,” in Speech and Computer , A. Karpov and R. Potapova, Eds. Springer International Publishing, 2020
2020
Later among the works it cites.
A. Łańcucki, “FastPitch: Parallel text-to-speech with pitch prediction,” arXiv:2006.06873 , 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
S. Majumdar, J. Balam, O. Hrinchuk, V. Lavrukhin, V. Noroozi, and B. Ginsburg, “Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition,” 2021
2021
Closest in time.