Fetching the paper…
Reading the bibliography…
This paper introduces a new speech corpus called "LibriTTS" designed for text-to-speech use.
J. L. Hintze and R. D. Nelson, “Violin plots: A box plot-density trace synergism,”
1998
Earlier work this paper cites.
J. Kominek and A. Black, “CMU ARCTIC databases for speech synthesis,” Carnegie Mellon University, Tech. Rep. CMU-LTI-03-177, 2003
2003
Earlier work this paper cites.
H. Zen, K. Tokuda, and A. Black, “Statistical parametric speech synthesis,”
2009
Earlier work this paper cites.
P. Taylor,
2009
Earlier work this paper cites.
C. Alberti and M. Bacchiani, “Automatic captioning in YouTube,”
2009
Earlier work this paper cites.
P. J. Moreno and C. Alberti, “A factor automaton approach for the forced alignment of long speech recordings,” in
2009
Earlier work this paper cites.
C. Kim and R. M. Stern, “Robust signal-to-noise ratio estimation based on waveform amplitude distribution analysis,” in
2009
Earlier work this paper cites.
S. King and V. Karaiskos, “The Blizzard Challenge 2011,” in
2011
Earlier work this paper cites.
H. Hofman, K. Kafadar, and H. Wickham, “Letter-value plots: Boxplots for large data,” had.co.nz, Tech. Rep., 2011
2011
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” University of Edinburgh. The Centre for Speech Technology Research (CSTR)
2012
Earlier work this paper cites.
——, “The Blizzard Challenge 2013,” in
2013
Earlier work this paper cites.
H. Liao, E. McDermott, and A. Senior, “Large scale deep neural network acoustic modeling with semi-supervised training data for YouTube video transcription,” in
2013
Earlier work this paper cites.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in
2015
Cited alongside, same era.
P. Ebden and R. Sproat, “The Kestrel TTS text normalization system,”
2015
Cited alongside, same era.
J. Sotelo, S. Mehri, K. Kumar, J. Santos, K. Kastner, A. Courville, and Y. Bengio, “Char2Wav: End-to-End speech synthesis,” in
2017
Cited alongside, same era.
S. Arik, C. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. Weiss, N. Jaitly, Z. Yang
2017
Cited alongside, same era.
S. Arik, G. Diamos, A. Gibiansky, J. Miller, K. Peng, W. Ping, J. Raiman
2017
2018
Later among the works it cites.
S. Arik, J. Chen, P. W. Peng, Kainan and, and Y. Zhou, “Neural voice cloning with a few samples,”
2018
Later among the works it cites.
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, A. Wang
2018
Later among the works it cites.
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan, E. Bettenberg, J. Shor, Y. Xiao
2018
Later among the works it cites.
W.-N. Hsu, Y. Zhang, R. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2017
Cited alongside, same era.
T. Capes, P. Coles, A. Conkie, L. Golipour, A. Hadjitarkhani, Q. Hu, N. Huddleston
2017
Cited alongside, same era.
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche
2017
Cited alongside, same era.
K. Ito, “The LJ speech dataset,”
2017
Cited alongside, same era.
J. Shen, R. Pang, J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen
2018
Cited alongside, same era.
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. Zhou, “Close to human quality TTS with transformer,”
2018
Cited alongside, same era.
Later among the works it cites.
“Open Data Commons Attribution License (ODC-By) v1.0,”
2018
Later among the works it cites.
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, “Fitting new speakers based on a short untranscribed sample,”
2018
Later among the works it cites.
Y. Lee, T. Kim, and S. Lee, “Voice imitating text-to-speech neural networks,”
2018
Later among the works it cites.
D.-R. Liu, C.-Y. Yang, S.-L. Wu, and H.-Y. Lee, “Improving unsupervised style transfer in end-to-end speech synthesis with end-to-end speech recognition,” in
2018
Later among the works it cites.
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. J. Weiss
2018
Later among the works it cites.
2018
Later among the works it cites.
K. Kastner, J. F. Santos, Y. Bengio, and A. C. Courville, “Representation mixing for TTS synthesis,”
2018
Later among the works it cites.
N. Kalchbrenner, E. Erich, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg
2018
Later among the works it cites.
Munich Artificial Intelligence Laboratories GmbH, “The M-AILABS speech dataset,”
2019
Closest in time.