Fetching the paper…
Reading the bibliography…
Recent neural Text-to-Speech (TTS) models have been shown to perform very well when enough data is available.
“From wer and ril to mer and wil: improved evaluation measures for connected speech recognition,”
A. Morris, V. Maier, and P. Green, · 2004
Earlier work this paper cites.
“Learning without forgetting,” 2017
Zhizhong Li and Derek Hoiem, · 2017
Earlier work this paper cites.
“The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito and Linda Johnson, · 2017
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” 2018
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu, · 2018
Earlier work this paper cites.
“Deep voice 3: Scaling text-to-speech with convolutional sequence learning,” 2018
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller, · 2018
Earlier work this paper cites.
“Semi-supervised training for improving data efficiency in end-to-end speech synthesis,” 2018
Yu-An Chung, Yuxuan Wang, Wei-Ning Hsu, Yu Zhang, and RJ Skerry-Ryan, · 2018
Earlier work this paper cites.
“Forward attention in sequence- to-sequence acoustic modeling for speech synthesis,”
Jing-Xuan Zhang, Zhen-Hua Ling, and Li-Rong Dai, · 2018
Cited alongside, same era.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” 2019
Ye Jia, Yu Zhang, Ron J. Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, and Yonghui Wu, · 2019
Cited alongside, same era.
“Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning,” 2019
Yu Zhang, Ron J. Weiss, Heiga Zen, Yonghui Wu, Zhifeng Chen, RJ Skerry-Ryan, Ye Jia, Andrew Rosenberg, and Bhuvana Ramabhadran, · 2019
Cited alongside, same era.
“Sample efficient adaptive text-to-speech,” 2019
Yutian Chen, Yannis Assael, Brendan Shillingford, David Budden, Scott Reed, Heiga Zen, Quan Wang, Luis C. Cobo, Andrew Trask, Ben Laurie, Caglar Gulcehre, Aäron van den Oord, Oriol Vinyals, and Nando de Freitas, · 2019
Cited alongside, same era.
“Cross-lingual multi-speaker text-to-speech synthesis for voice cloning without using parallel corpus for unseen speakers,” 2019
“Towards transfer learning for end-to-end speech synthesis from deep pre-trained language models,” 2019
Wei Fang, Yu-An Chung, and James Glass, · 2019
Later among the works it cites.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,” 2020
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2020
Closest in time.
“Boffin tts: Few-shot speaker adaptation by bayesian optimization,” 2020
Henry B. Moss, Vatsal Aggarwal, Nishant Prateek, Javier González, and Roberto Barra-Chicote, · 2020
Closest in time.
“End-to-end code-switching tts with cross-lingual language model,”
X. Zhou, X. Tian, G. Lee, R. K. Das, and H. Li, · 2020
Closest in time.
“An open source implementation of itu-t recommendation p. 808 with validation,”
Babak Naderi and Ross Cutler, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhaoyu Liu and Brian Mak, · 2019
Cited alongside, same era.
“Css10: A collection of single speaker speech datasets for 10 languages,”
Kyubyong Park and Thomas Mulc, · 2019
Cited alongside, same era.