Fetching the paper…
Reading the bibliography…
Recently, sequence-to-sequence models with attention have been successfully applied in Text-to-speech (TTS).
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in
1993
Earlier work this paper cites.
H. Lee, P. Pham, Y. Largman, and A. Y. Ng, “Unsupervised feature learning for audio classification using convolutional deep belief networks,” in
2009
Earlier work this paper cites.
O. S. Watts, “Unsupervised learning for text-to-speech synthesis,” Ph.D. dissertation, The University of Edinburgh, 2012
2012
Earlier work this paper cites.
J. Glass, “Towards unsupervised speech processing,” in
2012
Earlier work this paper cites.
N. Perraudin, P. Balazs, and P. L. Søndergaard, “A fast Griffin-Lim algorithm,” in
2013
Earlier work this paper cites.
O. Watts, Z. Wu, and S. King, “Sentence-level control vectors for deep neural network speech synthesis,” in
2015
Earlier work this paper cites.
P. Wang, Y. Qian, F. K. Soong, L. He, and H. Zhao, “Word embedding for recurrent neural network based tts synthesis,” in
2015
Earlier work this paper cites.
Q. Yu, P. Liu, Z. Wu, S. K. Ang, H. Meng, and L. Cai, “Learning cross-lingual information with multilingual BLSTM for speech synthesis of low-resource languages,” in
2016
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Earlier work this paper cites.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. C. Courville, and Y. Bengio, “Char2wav: End-to-end speech synthesis,” in
2017
Earlier work this paper cites.
E. Cooper and X. Wang, “Utterance selection for optimizing intelligibility of TTS voices trained on asr data,”
2017
Earlier work this paper cites.
A. Gutkin, “Uniform multilingual multi-speaker acoustic model for statistical parametric speech synthesis of low-resourced languages,”
2017
Earlier work this paper cites.
W.-N. Hsu, Y. Zhang, and J. Glass, “Unsupervised learning of disentangled and interpretable representations from sequential data,” in
2017
Cited alongside, same era.
A. Van Den Oord, O. Vinyals
2017
Cited alongside, same era.
K. Ito, “The LJ speech dataset,”
2017
Cited alongside, same era.
C. Veaux, J. Yamagishi, and K. MacDonald, “Superseded - CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2017. [Online]. Available:
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu
2018
Later among the works it cites.
F.-Y. Kuo, S. Aryal, G. Degottex, S. Kang, P. Lanchantin, and I. Ouyang, “Data selection for improving naturalness of TTS voices trained on small found corpuses,” in
2018
Later among the works it cites.
J. Park, K. Zhao, K. Peng, and W. Ping, “Multi-speaker end-to-end speech synthesis,”
2019
Later among the works it cites.
F. Kuo, I. Ouyang, S. Aryal, and P. Lanchantin, “Selection and training schemes for improving TTS voice built on found data,”
2019
Later among the works it cites.
J. Fong, P. O. Gallegos, Z. Hodari, and S. King, “Investigating the robustness of sequence-to-sequence text-to-speech models to imperfectly-transcribed training data,” in
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep Voice 3: Scaling text-to-speech with convolutional sequence learning,” in
2018
Cited alongside, same era.
W. Ping, K. Peng, and J. Chen, “Clarinet: Parallel wave generation in end-to-end text-to-speech,” in
2018
Cited alongside, same era.
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, Q. Wang, L. C. Cobo, A. Trask, B. Laurie
2018
Cited alongside, same era.
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, “Fitting new speakers based on a short untranscribed sample,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in
2018
Cited alongside, same era.
Later among the works it cites.
Y.-A. Chung, Y. Wang, W.-N. Hsu, Y. Zhang, and R. Skerry-Ryan, “Semi-supervised training for improving data efficiency in end-to-end speech synthesis,” in
2019
Later among the works it cites.
Y.-J. Chen, T. Tu, C.-c. Yeh, and H.-Y. Lee, “End-to-end text-to-speech for low-resource languages by cross-lingual transfer learning,”
2019
Later among the works it cites.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An unsupervised autoregressive model for speech representation learning,”
2019
Later among the works it cites.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using WaveNet autoencoders,”
2019
Later among the works it cites.
H. B. Moss, V. Aggarwal, N. Prateek, J. González, and R. Barra-Chicote, “BOFFIN TTS: Few-shot speaker adaptation by bayesian optimization,” in
2020
Closest in time.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common Voice: A massively-multilingual speech corpus,” in
2020
Closest in time.