Fetching the paper…
Reading the bibliography…
Recently, there has been a growing interest in the field of controllable Text-to-Speech (TTS).
“Toronto emotional speech set (tess)-younger talker_happy,”
Kate Dupuis and M Kathleen Pichora-Fuller, · 2010
Earlier work this paper cites.
“World: a vocoder-based high-quality speech synthesis system for real-time applications,”
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa, · 2016
Earlier work this paper cites.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,”
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al., · 2017
Earlier work this paper cites.
“Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,”
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A Saurous, · 2018
Earlier work this paper cites.
“Libritts: A corpus derived from librispeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Earlier work this paper cites.
“Categorical and dimensional ratings of emotional speech: Behavioral findings from the morgan emotional speech set,”
Shae D Morgan, · 2019
Earlier work this paper cites.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2020
Earlier work this paper cites.
“Speaking speed control of end-to-end speech synthesis using sentence-level conditioning,”
Jae-Sung Bae, Hanbin Bae, Young-Sun Joo, Junmo Lee, Gyeong-Hoon Lee, and Hoon-Young Cho, · 2020
Cited alongside, same era.
“Mead: A large-scale audio-visual dataset for emotional talking-face generation,”
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy, · 2020
Cited alongside, same era.
“Fastpitchformant: Source-filter based decomposed modeling for speech synthesis,”
Taejun Bak, Jae-Sung Bae, Hanbin Bae, Young-Ik Kim, and Hoon-Young Cho, · 2021
Cited alongside, same era.
“Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,”
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li, · 2021
Cited alongside, same era.
“Prompt programming for large language models: Beyond the few-shot paradigm,”
Laria Reynolds and Kyle McDonell, · 2021
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Later among the works it cites.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Closest in time.
“Mega-tts: Zero-shot text-to-speech at scale with intrinsic inductive bias,”
Ziyue Jiang, Yi Ren, Zhenhui Ye, Jinglin Liu, Chen Zhang, Qian Yang, Shengpeng Ji, Rongjie Huang, Chunfeng Wang, Xiang Yin, et al., · 2023
Closest in time.
“Prompttts: Controllable text-to-speech with text descriptions,”
Zhifang Guo, Yichong Leng, Yihan Wu, Sheng Zhao, and Xu Tan, · 2023
Closest in time.
“Promptstyle: Controllable style transfer for text-to-speech with natural language descriptions,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Clipcap: Clip prefix for image captioning,”
Ron Mokady, Amir Hertz, and Amit H Bermano, · 2021
Cited alongside, same era.
Dongchao Yang, Songxiang Liu, Jianwei Yu, Helin Wang, Chao Weng, and Yuexian Zou, · 2022
Cited alongside, same era.
Guanghou Liu, Yongmao Zhang, Yi Lei, Yunlin Chen, Rui Wang, Zhifei Li, and Lei Xie, · 2023
Closest in time.
“Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,”
Dongchao Yang, Songxiang Liu, Rongjie Huang, Guangzhi Lei, Chao Weng, Helen Meng, and Dong Yu, · 2023
Closest in time.