Fetching the paper…
Reading the bibliography…
Recent advancements in personalized speech generation have brought synthetic speech increasingly close to the realism of target speakers' recordings, yet multimodal speaker generation remains on the rise.
Facenet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
LRS3-TED: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
GPU accelerated t-distributed stochastic neighbor embedding
David M. Chan, Roshan Rao, Forrest Huang, and John F. Canny · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Earlier work this paper cites.
Video prediction recalling long-term motion context via memory alignment learning
Sangmin Lee, Hak Gu Kim, Dae Hwi Choi, Hyung-Il Kim, and Yong Man Ro · 2021
Earlier work this paper cites.
Face-based voice conversion: Learning the voice behind a face
Hsiao-Han Lu, Shao-En Weng, Ya-Fan Yen, Hong-Han Shuai, and Wen-Huang Cheng · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Naturalspeech: End-to-end text to speech synthesis with human-level quality
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, Frank K. Soong, Tao Qin, Sheng Zhao, and Tie-Yan Liu · 2022
Earlier work this paper cites.
Prompttts: Controllable text-to-speech with text descriptions
Zhifang Guo, Yichong Leng, Yihan Wu, Sheng Zhao, and Xu Tan · 2023
Earlier work this paper cites.
Minki Kang, Wooseok Han, and Eunho Yang · 2023
Earlier work this paper cites.
Imaginary voice: Face-styled diffusion model for text-to-speech
Jiyoung Lee, Joon Son Chung, and Soo-Whan Chung · 2023
Earlier work this paper cites.
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi · 2023
Earlier work this paper cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Earlier work this paper cites.
Face-driven zero-shot voice conversion with memory-based face-voice alignment
Zhengyan Sheng, Yang Ai, Yan-Nian Chen, and Zhen-Hua Ling · 2023
Cited alongside, same era.
Conditional flow matching: Simulation-free dynamic optimal transport
Alexander Tong, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Kilian Fatras, Guy Wolf, and Yoshua Bengio · 2023
Cited alongside, same era.
Audiobox: Unified audio generation with natural language prompts
Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra, Yi-Chiao Wu, Baishan Guo, Jiemin Zhang, Xinyue Zhang, Robert Adkins, William Ngan, Jeff Wang, Ivan Cruz, Bapi Akula, Akinniyi Akinyemi, Brian Ellis, Rashel Moritz, Yael Yungster, Alice Rakotoarison, Liang Tan, Chris Summers, Carleigh Wood, Joshua Lane, Mary Williamson, and Wei-Ning Hsu · 2023
Cited alongside, same era.
Neural codec language models are zero-shot text to speech synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei · 2023
Cited alongside, same era.
Textrolspeech: A text style control speech corpus with codec language text-to-speech models
Shengpeng Ji, Jialong Zuo, Minghui Fang, Ziyue Jiang, Feiyang Chen, Xinyu Duan, Baoxing Huai, and Zhou Zhao · 2024
Later among the works it cites.
Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Eric Liu, Yichong Leng, Kaitao Song, Siliang Tang, Zhizheng Wu, Tao Qin, Xiangyang Li, Wei Ye, Shikun Zhang, Jiang Bian, Lei He, Jinyu Li, and Sheng Zhao · 2024
Later among the works it cites.
Masaya Kawamura, Ryuichi Yamamoto, Yuma Shirahata, Takuya Hasumi, and Kentaro Tachibana · 2024
Later among the works it cites.
Clam-tts: Improving neural codec language model for zero-shot text-to-speech
Jaehyeon Kim, Keon Lee, Seungjun Chung, and Jaewoong Cho · 2024
Later among the works it cites.
Ditto-tts: Efficient and scalable zero-shot text-to-speech with diffusion transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CAM++: A fast and efficient network for speaker verification using context-aware masking
Hui Wang, Siqi Zheng, Yafeng Chen, Luyao Cheng, and Qian Chen · 2023
Cited alongside, same era.
Zero-shot face-based voice conversion: Bottleneck-free speech disentanglement in the real-world scenario
Shao-En Weng, Hong-Han Shuai, and Wen-Huang Cheng · 2023
Cited alongside, same era.
Promptspeaker: Speaker generation based on text descriptions
Yongmao Zhang, Guanghou Liu, Yi Lei, Yunlin Chen, Hao Yin, Lei Xie, and Zhifei Li · 2023
Cited alongside, same era.
Seed-tts: A family of high-quality versatile speech generation models
Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, Mingqing Gong, Peisong Huang, Qingqing Huang, Zhiying Huang, Yuanyuan Huo, Dongya Jia, Chumin Li, Feiya Li, Hui Li, Jiaxin Li, Xiaoyang Li, Xingxing Li, Lin Liu, Shouda Liu, Sichao Liu, Xudong Liu, Yuchen Liu, Zhengxi Liu, Lu Lu, Junjie Pan, Xin Wang, Yuping Wang, Yuxuan Wang, Zhen Wei, Jian Wu, Chao Yao, Yifeng Yang, Yuanhao Yi, Junteng Zhang, Qidi Zhang, Shuo Zhang, Wenjie Zhang, Yang Zhang, Zilin Zhao, Dejian Zhong, and Xiaobin Zhuang · 2024
Cited alongside, same era.
VALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers
Sanyuan Chen, Shujie Liu, Long Zhou, Yanqing Liu, Xu Tan, Jinyu Li, Sheng Zhao, Yao Qian, and Furu Wei · 2024
Cited alongside, same era.
Zhihao Du, Qian Chen, Shiliang Zhang, Kai Hu, Heng Lu, Yexin Yang, Hangrui Hu, Siqi Zheng, Yue Gu, Ziyang Ma, Zhifu Gao, and Zhijie Yan · 2024
Cited alongside, same era.
Softclip: Softer cross-modal alignment makes CLIP stronger
Yuting Gao, Jinfeng Liu, Zihan Xu, Tong Wu, Enwei Zhang, Ke Li, Jie Yang, Wei Liu, and Xing Sun · 2024
Cited alongside, same era.
MM-TTS: multi-modal prompt based style transfer for expressive text-to-speech synthesis
Wenhao Guan, Yishuang Li, Tao Li, Hukai Huang, Feng Wang, Jiayan Lin, Lingyan Huang, Lin Li, and Qingyang Hong · 2024
Cited alongside, same era.
Keon Lee, Dong Won Kim, Jaehyeon Kim, and Jaewoong Cho · 2024
Later among the works it cites.
Prompttts 2: Describing and generating voices with text prompt
Yichong Leng, Zhifang Guo, Kai Shen, Zeqian Ju, Xu Tan, Eric Liu, Yufei Liu, Dongchao Yang, Leying Zhang, Kaitao Song, Lei He, Xiangyang Li, Sheng Zhao, Tao Qin, and Jiang Bian · 2024
Later among the works it cites.
SYNTHE-SEES: face based text-to-speech for virtual speaker
Jae Hyun Park, Joon-Gyu Maeng, Taejun Bak, and Young-Sun Joo · 2024
Later among the works it cites.
Voice attribute editing with text prompt
Zhengyan Sheng, Yang Ai, Li-Juan Liu, Jia Pan, and Zhen-Hua Ling · 2024
Later among the works it cites.
Prompttts++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions
Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, Hironori Doi, Tatsuya Komatsu, and Kentaro Tachibana · 2024
Later among the works it cites.
Multiscale matching driven by cross-modal similarity consistency for audio-text retrieval
Qian Wang, Jia-Chen Gu, and Zhen-Hua Ling · 2024
Later among the works it cites.
Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt
Dongchao Yang, Songxiang Liu, Rongjie Huang, Chao Weng, and Helen Meng · 2024
Later among the works it cites.
Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts
Jixun Yao, Yuguang Yang, Yi Lei, Ziqian Ning, Yanni Hu, Yu Pan, Jingjing Yin, Hongbin Zhou, Heng Lu, and Lei Xie · 2024
Later among the works it cites.
ASQ: an ultra-low bit rate asr-oriented speech quantization method
Lingxuan Ye, Changfeng Gao, Gaofeng Cheng, Liuping Luo, and Qingwei Zhao · 2024
Later among the works it cites.