Fetching the paper…
Reading the bibliography…
Despite recent advancements in speech generation with text prompt providing control over speech style, voice attributes in synthesized speech remain elusive and challenging to control.
Jason Weston, Sumit Chopra, and Antoine Bordes. 2015 · 2015
Earlier work this paper cites.
An overview of voice conversion systems
Seyed Hamidreza Mohammadi and Alexander Kain. 2017 · 2017
Earlier work this paper cites.
Describing sound: The cognitive linguistics of timbre
Zachary Wallmark and Roger A Kendall. 2018 · 2018
Earlier work this paper cites.
GPU accelerated t-distributed stochastic neighbor embedding
David M. Chan, Roshan Rao, Forrest Huang, and John F. Canny. 2019 · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2020
Earlier work this paper cites.
Cross-modal memory networks for radiology report generation
Zhihong Chen, Yaling Shen, Yan Song, and Xiang Wan. 2021 · 2021
Cited alongside, same era.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Content-dependent fine-grained speaker embedding for zero-shot speaker adaptation in text-to-speech synthesis
Yixuan Zhou, Changhe Song, Xiang Li, Luwen Zhang, Zhiyong Wu, Yanyao Bian, Dan Su, and Helen Meng. 2022 · 2022
Cited alongside, same era.
Prompttts: Controllable text-to-speech with text descriptions
Zhifang Guo, Yichong Leng, Yihan Wu, Sheng Zhao, and Xu Tan. 2023 · 2023
Cited alongside, same era.
Textrolspeech: A text style control speech corpus with codec language text-to-speech models
Shengpeng Ji, Jialong Zuo, Minghui Fang, Ziyue Jiang, Feiyang Chen, Xinyu Duan, Baoxing Huai, and Zhou Zhao. 2023 · 2023
Promptstyle: Controllable style transfer for text-to-speech with natural language descriptions
Guanghou Liu, Yongmao Zhang, Yi Lei, Yunlin Chen, Rui Wang, Zhifei Li, and Lei Xie. 2023 · 2023
Later among the works it cites.
Face-driven zero-shot voice conversion with memory-based face-voice alignment
Zhengyan Sheng, Yang Ai, Yan-Nian Chen, and Zhen-Hua Ling. 2023 · 2023
Later among the works it cites.
Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, Hironori Doi, Tatsuya Komatsu, and Kentaro Tachibana. 2023 · 2023
Later among the works it cites.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Christophe Veaux, Junichi Yamagishi, and Kirsten MacDonald. 2023 · 2023
Later among the works it cites.
COCO-NUT: corpus of japanese utterance and voice characteristics description for prompt-based control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Imagic: Text-based real image editing with diffusion models
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. 2023 · 2023
Cited alongside, same era.
Prompttts 2: Describing and generating voices with text prompt
Yichong Leng, Zhifang Guo, Kai Shen, Xu Tan, Zeqian Ju, Yanqing Liu, Yufei Liu, Dongchao Yang, Leying Zhang, Kaitao Song, Lei He, Xiang-Yang Li, Sheng Zhao, Tao Qin, and Jiang Bian. 2023 · 2023
Cited alongside, same era.
Freevc: Towards high-quality text-free one-shot voice conversion
Jingyi Li, Weiping Tu, and Li Xiao. 2023 · 2023
Cited alongside, same era.
Aya Watanabe, Shinnosuke Takamichi, Yuki Saito, Wataru Nakata, Detai Xin, and Hiroshi Saruwatari. 2023 · 2023
Later among the works it cites.
Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt
Dongchao Yang, Songxiang Liu, Rongjie Huang, Guangzhi Lei, Chao Weng, Helen Meng, and Dong Yu. 2023 · 2023
Later among the works it cites.
Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts
Jixun Yao, Yuguang Yang, Yi Lei, Ziqian Ning, Yanni Hu, Yu Pan, Jingjing Yin, Hongbin Zhou, Heng Lu, and Lei Xie. 2023 · 2023
Later among the works it cites.
Promptspeaker: Speaker generation based on text descriptions
Yongmao Zhang, Guanghou Liu, Yi Lei, Yunlin Chen, Hao Yin, Lei Xie, and Zhifei Li. 2023 · 2023
Later among the works it cites.