Fetching the paper…
Reading the bibliography…
Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair.
“The jensen-shannon divergence,”
María Luisa Menéndez, JA Pardo, L Pardo, and MC Pardo, · 1997
Earlier work this paper cites.
“A style control technique for hmm-based expressive speech synthesis,”
Takashi Nose, Junichi Yamagishi, Takashi Masuko, and Takao Kobayashi, · 2007
Earlier work this paper cites.
“Emotional end-to-end neural speech synthesizer,”
Younggun Lee, Azam Rabiee, and Soo-Young Lee, · 2017
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A Saurous, · 2018
Earlier work this paper cites.
“Emotional speech synthesis with rich and granularized control,”
Se-Yun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, ChungHyun Ahn, and Hong-Goo Kang, · 2020
Earlier work this paper cites.
“Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis,”
Yi Lei, Shan Yang, and Lei Xie, · 2021
Earlier work this paper cites.
“Controllable emotion transfer for end-to-end speech synthesis,”
Tao Li, Shan Yang, Liumeng Xue, and Lei Xie, · 2021
Earlier work this paper cites.
“Expressive text-to-speech using style tag,”
Minchan Kim, Sung Jun Cheon, Byoung Jin Choi, Jong Jin Kim, and Nam Soo Kim, · 2021
Earlier work this paper cites.
“Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,”
Xiong Cai, Dongyang Dai, Zhiyong Wu, Xiang Li, Jingbei Li, and Helen Meng, · 2021
Earlier work this paper cites.
“Text to speech synthesis: A systematic review, deep learning based architecture and future research direction,”
Fahima Khanam, Farha Akhter Munmun, Nadia Afrin Ritu, Aloke Kumar Saha, and Muhammad Firoz, · 2022
Earlier work this paper cites.
“Speech synthesis with mixed emotions,”
Kun Zhou, Berrak Sisman, Rajib Rana, Björn W Schuller, and Haizhou Li, · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback,”
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, et al., · 2022
Earlier work this paper cites.
“Emotional voice conversion: Theory, databases and esd,”
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li, · 2022
Cited alongside, same era.
“Text-to-speech synthesis based on latent variable conversion using diffusion probabilistic model and variational autoencoder,”
Yusuke Yasuda and Tomoki Toda, · 2023
Cited alongside, same era.
“A vector quantized approach for text to speech synthesis on real-world spontaneous speech,”
Li-Wei Chen, Shinji Watanabe, and Alexander Rudnicky, · 2023
Cited alongside, same era.
“Emospeech: guiding fastspeech2 towards emotional text to speech,”
Daria Diatlova and Vitalii Shutov, · 2023
Cited alongside, same era.
“An emotion speech synthesis method based on vits,”
Wei Zhao and Zheng Yang, · 2023
Cited alongside, same era.
“Mm-tts: A unified framework for multimodal, prompt-induced emotional text-to-speech synthesis,”
Xiang Li, Zhi-Qi Cheng, Jun-Yan He, Xiaojiang Peng, and Alexander G Hauptmann, · 2024
Closest in time.
“Emotion rendering for conversational speech synthesis with heterogeneous graph-based context modeling,”
Rui Liu, Yifan Hu, Yi Ren, Xiang Yin, and Haizhou Li, · 2024
Closest in time.
“Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,”
Dongchao Yang, Songxiang Liu, Rongjie Huang, Chao Weng, and Helen Meng, · 2024
Closest in time.
Haibin Wu, Xiaofei Wang, Sefik Emre Eskimez, Manthan Thakker, Daniel Tompkins, Chung-Hsien Tsai, Canrun Li, Zhen Xiao, Sheng Zhao, Jinyu Li, et al., · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yiwei Guo, Chenpeng Du, Xie Chen, and Kai Yu, · 2023
Cited alongside, same era.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Cited alongside, same era.
“Direct preference optimization: Your language model is secretly a reward model,”
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn, · 2023
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al., · 2023
Cited alongside, same era.
“Gemini: a family of highly capable multimodal models,”
Gemini Team Google, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al., · 2023
Cited alongside, same era.
“Seamless: Multilingual expressive and streaming speech translation,”
Loïc Barrault, Yu-An Chung, Mariano Coria Meglioli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, et al., · 2023
Cited alongside, same era.
“emotion2vec: Self-supervised pre-training for speech emotion representation,”
Ziyang Ma, Zhisheng Zheng, Jiaxin Ye, Jinchao Li, Zhifu Gao, Shiliang Zhang, and Xie Chen, · 2023
Cited alongside, same era.
Zhihao Du, Qian Chen, Shiliang Zhang, Kai Hu, Heng Lu, Yexin Yang, Hangrui Hu, Siqi Zheng, Yue Gu, Ziyang Ma, et al., · 2024
Closest in time.
“Towards a unified view of preference learning for large language models: A survey,”
Bofei Gao, Feifan Song, Yibo Miao, Zefan Cai, Zhe Yang, Liang Chen, Helan Hu, Runxin Xu, Qingxiu Dong, Ce Zheng, Wen Xiao, Ge Zhang, Daoguang Zan, Keming Lu, Bowen Yu, Dayiheng Liu, Zeyu Cui, Jian Yang, Lei Sha, Houfeng Wang, Zhifang Sui, Peiyi Wang, Tianyu Liu, and Baobao Chang, · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al., · 2024
Closest in time.
“Speechalign: Aligning speech generation to human preferences,”
Dong Zhang, Zhaowei Li, Shimin Li, Xin Zhang, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu, · 2024
Closest in time.
“Musicrl: Aligning music generation to human preferences,”
Geoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent, Matej Kastelic, Zalán Borsos, Brian McWilliams, Victor Ungureanu, Olivier Bachem, Olivier Pietquin, et al., · 2024
Closest in time.
“Boost your own human image generation model via direct preference optimization with ai feedback,”
Sanghyeon Na, Yonggyu Kim, and Hyunjoon Lee, · 2024
Closest in time.
“Diffusion model alignment using direct preference optimization,”
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik, · 2024
Closest in time.