Fetching the paper…
Reading the bibliography…
There has been a significant progress in Text-To-Speech (TTS) synthesis technology in recent years, thanks to the advancement in neural generative modeling.
“Reverse-time diffusion equation models,”
Brian D.O. Anderson, · 1982
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Neural voice cloning with a few samples,”
Sercan Ömer Arik, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou, · 2018
Earlier work this paper cites.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Ye Jia, Yu Zhang, Ron J. Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez-Moreno, and Yonghui Wu, · 2018
Earlier work this paper cites.
“Sample efficient adaptive text-to-speech,”
Yutian Chen, Yannis M. Assael, Brendan Shillingford, David Budden, Scott E. Reed, Heiga Zen, Quan Wang, Luis C. Cobo, Andrew Trask, Ben Laurie, Çaglar Gülçehre, Aäron van den Oord, Oriol Vinyals, and Nando de Freitas, · 2019
Earlier work this paper cites.
“Libritts: A corpus derived from librispeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Earlier work this paper cites.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),”
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald, · 2019
Earlier work this paper cites.
“Glow-tts: A generative flow for text-to-speech via monotonic alignment search,”
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon, · 2020
Earlier work this paper cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Cited alongside, same era.
“Multispeech: Multi-speaker text to speech with transformer,”
Mingjian Chen, Xu Tan, Yi Ren, Jin Xu, Hao Sun, Sheng Zhao, and Tao Qin, · 2020
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2021
Cited alongside, same era.
“Score-based generative modeling through stochastic differential equations,”
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole, · 2021
Cited alongside, same era.
“Diff-tts: A denoising diffusion model for text-to-speech,”
“Meta-stylespeech : Multi-speaker adaptive text-to-speech generation,”
Dongchan Min, Dong Bok Lee, Eunho Yang, and Sung Ju Hwang, · 2021
Later among the works it cites.
“Adaspeech: Adaptive text to speech for custom voice,”
Mingjian Chen, Xu Tan, Bohan Li, Yanqing Liu, Tao Qin, Sheng Zhao, and Tie-Yan Liu, · 2021
Later among the works it cites.
“Guided-tts: A diffusion model for text-to-speech via classifier guidance,”
Heeseung Kim, Sungwon Kim, and Sungroh Yoon, · 2022
Closest in time.
“Yourtts: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,”
Edresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior, Eren Gölge, and Moacir A. Ponti, · 2022
Closest in time.
“Revisiting over-smoothness in text to speech,”
Yi Ren, Xu Tan, Tao Qin, Zhou Zhao, and Tie-Yan Liu, · 2022
Closest in time.
“One TTS alignment to rule them all,”
Rohan Badlani, Adrian Lancucki, Kevin J. Shih, Rafael Valle, Wei Ping, and Bryan Catanzaro, · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Myeonghun Jeong, Hyeongju Kim, Sung Jun Cheon, Byoung Jin Choi, and Nam Soo Kim, · 2021
Cited alongside, same era.
“Grad-tts: A diffusion probabilistic model for text-to-speech,”
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail A. Kudinov, · 2021
Cited alongside, same era.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
Jaehyeon Kim, Jungil Kong, and Juhee Son, · 2021
Cited alongside, same era.
Closest in time.
“Diffusion-based voice conversion with fast maximum likelihood sampling scheme,”
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, Mikhail Sergeevich Kudinov, and Jiansheng Wei, · 2022
Closest in time.
“Emotional voice conversion: Theory, databases and ESD,”
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li, · 2022
Closest in time.