Fetching the paper…
Reading the bibliography…
Style voice conversion aims to transform the style of source speech to a desired style according to real-world application demands.
“Mel-cepstral distance measure for objective speech quality assessment,”
Robert Kubichek, · 1993
Earlier work this paper cites.
“Pearson correlation coefficient,”
Israel Cohen, Yiteng Huang, Jingdong Chen, Jacob Benesty, Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen, · 2009
Earlier work this paper cites.
“A short-time objective intelligibility measure for time-frequency weighted noisy speech,”
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen, · 2010
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, R. J. Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A. Saurous, · 2018
Earlier work this paper cites.
“Unsupervised end-to-end learning of discrete linguistic units for voice conversion,”
Andy T. Liu, Po-Chun Hsu, and Hung-yi Lee, · 2019
Earlier work this paper cites.
“Sequence-to-sequence modelling of F0 for speech emotion conversion,”
Carl Robinson, Nicolas Obin, and Axel Roebel, · 2019
Earlier work this paper cites.
“Fastspeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Earlier work this paper cites.
“An overview of voice conversion and its challenges: From statistical modeling to deep learning,”
Berrak Sisman, Junichi Yamagishi, Simon King, and Haizhou Li, · 2020
Earlier work this paper cites.
“Converting anyone’s emotion: Towards speaker-independent emotional voice conversion,”
Kun Zhou, Berrak Sisman, Mingyang Zhang, and Haizhou Li, · 2020
Earlier work this paper cites.
“Seen and unseen emotional style transfer for voice conversion with A new emotional speech dataset,”
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li, · 2021
Cited alongside, same era.
“Expressive voice conversion: A joint framework for speaker identity and emotional style transfer,”
Zongyang Du, Berrak Sisman, Kun Zhou, and Haizhou Li, · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Meta-stylespeech : Multi-speaker adaptive text-to-speech generation,”
Dongchan Min, Dong Bok Lee, Eunho Yang, and Sung Ju Hwang, · 2021
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2021
Cited alongside, same era.
“One-shot voice conversion for style transfer based on speaker adaptation,”
Zhichao Wang, Qicong Xie, Tao Li, Hongqiang Du, Lei Xie, Pengcheng Zhu, and Mengxiao Bi, · 2022
Later among the works it cites.
“High-resolution image synthesis with latent diffusion models,”
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer, · 2022
Later among the works it cites.
“Mixed emotion modelling for emotional voice conversion,”
Kun Zhou, Berrak Sisman, Carlos Busso, and Haizhou Li, · 2022
Later among the works it cites.
“Preserving background sound in noise-robust voice conversion via multi-task learning,”
Jixun Yao, Yi Lei, Qing Wang, Pengcheng Guo, Ziqian Ning, Lei Xie, Hai Li, Junhui Liu, and Danming Xie, · 2023
Closest in time.
“Distinguishable speaker anonymization based on formant and fundamental frequency scaling,”
Jixun Yao, Qing Wang, Yi Lei, Pengcheng Guo, Lei Xie, Namin Wang, and Jie Liu, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
Jaehyeon Kim, Jungil Kong, and Juhee Son, · 2021
Cited alongside, same era.
“Self-supervised context-aware style representation for expressive speech synthesis,”
Yihan Wu, Xi Wang, Shaofei Zhang, Lei He, Ruihua Song, and Jian-Yun Nie, · 2022
Cited alongside, same era.
“Emotional voice conversion: Theory, databases and ESD,”
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li, · 2022
Cited alongside, same era.
“Cross-speaker style transfer for text-to-speech using data augmentation,”
Manuel Sam Ribeiro, Julian Roth, Giulia Comini, Goeric Huybrechts, Adam Gabrys, and Jaime Lorenzo-Trueba, · 2022
Cited alongside, same era.
“Nonparallel emotional voice conversion for unseen speaker-emotion pairs using dual domain adversarial network virtual domain pairing,”
Nirmesh Shah, Mayank Singh, Naoya Takahashi, and Naoyuki Onoe, · 2023
Closest in time.
“InstructTTS: Modelling expressive tts in discrete latent space with natural language style prompt,”
Dongchao Yang, Songxiang Liu, Rongjie Huang, Guangzhi Lei, Chao Weng, Helen Meng, and Dong Yu, · 2023
Closest in time.
“PromptStyle: Controllable style transfer for text-to-speech with natural language descriptions,”
Guanghou Liu, Yongmao Zhang, Yi Lei, Yunlin Chen, Rui Wang, Zhifei Li, and Lei Xie, · 2023
Closest in time.
“PromptTTS: Controllable text-to-speech with text descriptions,”
Zhifang Guo, Yichong Leng, Yihan Wu, Sheng Zhao, and Xu Tan, · 2023
Closest in time.