Fetching the paper…
Reading the bibliography…
We work to create a multilingual speech synthesis system which can generate speech with the proper accent while retaining the characteristics of an individual voice.
“Praat: doing phonetics by computer,”
Paul Boersma and David Weenink, · 2003
Earlier work this paper cites.
“Multi-language multi-speaker acoustic modeling for lstm-rnn based statistical parametric speech synthesis,”
Bo Li and Heiga Zen, · 2016
Earlier work this paper cites.
“Tacotron: A fully end-to-end text-to-speech synthesis model,”
Yuxuan Wang, R. J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc V. Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous, · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, R. J. Skerry-Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu, · 2017
Earlier work this paper cites.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Ye Jia, Yu Zhang, Ron J. Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez-Moreno, and Yonghui Wu, · 2018
Earlier work this paper cites.
“Unsupervised polyglot text-to-speech,”
Eliya Nachmani and Lior Wolf, · 2019
Earlier work this paper cites.
Yu Zhang, Ron J. Weiss, Heiga Zen, Yonghui Wu, Zhifeng Chen, R. J. Skerry-Ryan, Ye Jia, Andrew Rosenberg, and Bhuvana Ramabhadran, · 2019
Earlier work this paper cites.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Cited alongside, same era.
“Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis,”
Rafael Valle, Kevin J Shih, Ryan Prenger, and Bryan Catanzaro, · 2020
Cited alongside, same era.
“Fastpitch: Parallel text-to-speech with pitch prediction,” 2020
Adrian Łańcucki, · 2020
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,” 2020
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang, · 2020
Cited alongside, same era.
“Yourtts: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,”
Edresson Casanova, Julian Weber, Christopher Shulby, Arnaldo Cândido Júnior, Eren Gölge, and Moacir Antonelli Ponti, · 2021
Later among the works it cites.
“Titanet: Neural model for speaker representation with 1d depth-wise separable convolutions and global context,” 2021
Nithin Rao Koluguri, Taejin Park, and Boris Ginsburg, · 2021
Later among the works it cites.
“Naturalspeech: End-to-end text to speech synthesis with human-level quality,” 2022
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, Frank Soong, Tao Qin, Sheng Zhao, and Tie-Yan Liu, · 2022
Later among the works it cites.
“Generative modeling for low dimensional speech attributes with neural spline flows,” 2022
Kevin J. Shih, Rafael Valle, Rohan Badlani, João Felipe Santos, and Bryan Catanzaro, · 2022
Later among the works it cites.
“One tts alignment to rule them all,”
Rohan Badlani, Adrian Łańcucki, Kevin J. Shih, Rafael Valle, Wei Ping, and Bryan Catanzaro, · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“RAD-TTS: Parallel flow-based TTS with robust alignment learning and diverse synthesis,”
Kevin J. Shih, Rafael Valle, Rohan Badlani, Adrian Lancucki, Wei Ping, and Bryan Catanzaro, · 2021
Cited alongside, same era.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
Jaehyeon Kim, Jungil Kong, and Juhee Son, · 2021
Cited alongside, same era.
Later among the works it cites.
“VICReg: Variance-invariance-covariance regularization for self-supervised learning,”
Adrien Bardes, Jean Ponce, and Yann LeCun, · 2022
Later among the works it cites.
“Revisiting over-smoothness in text to speech,”
Yi Ren, Xu Tan, Tao Qin, Zhou Zhao, and Tie-Yan Liu, · 2022
Later among the works it cites.