Fetching the paper…
Reading the bibliography…
Zero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content.
Dynamic time warping
Müller, M. 2007 · 2007
Earlier work this paper cites.
Query-key normalization for transformers
Henry, A.; Dachapally, P. R.; Pawar, S.; and Chen, Y. 2020 · 2010
Earlier work this paper cites.
Pearson’s correlation coefficient
Sedgwick, P. 2012 · 2012
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Ganin, Y.; and Lempitsky, V. 2015 · 2015
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Perez, E.; Strub, F.; De Vries, H.; Dumoulin, V.; and Courville, A. 2018 · 2018
Earlier work this paper cites.
Autovc: Zero-shot voice style transfer with only autoencoder loss
Qian, K.; Zhang, Y.; Chang, S.; Yang, X.; and Hasegawa-Johnson, M. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
CSTR VCTK corpus: English multi-speaker corpus for cstr voice cloning toolkit
Veaux, C.; Yamagishi, J.; MacDonald, K.; et al. 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Libri-light: A benchmark for asr with limited or no supervision
Kahn, J.; Riviere, M.; Zheng, W.; Kharitonov, E.; Xu, Q.; Mazaré, P.-E.; Karadayi, J.; Liptchinsky, V.; Collobert, R.; Fuegen, C.; et al. 2020 · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J.; Kim, J.; and Bae, J. 2020 · 2020
Earlier work this paper cites.
Diffusion-based voice conversion with fast maximum likelihood sampling scheme
Popov, V.; Vovk, I.; Gogoryan, V.; Sadekova, T.; Kudinov, M.; and Wei, J. 2021 · 2021
Earlier work this paper cites.
Wang, D.; Deng, L.; Yeung, Y. T.; Chen, X.; Liu, X.; and Meng, H. 2021 · 2021
Earlier work this paper cites.
Improving zero-shot voice style transfer via disentangled representation learning
Yuan, S.; Cheng, P.; Zhang, R.; Hao, W.; Gan, Z.; and Carin, L. 2021 · 2021
Earlier work this paper cites.
Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone
Casanova, E.; Weber, J.; Shulby, C. D.; Junior, A. C.; Gölge, E.; and Ponti, M. A. 2022 · 2022
Earlier work this paper cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Chen, S.; Wang, C.; Chen, Z.; Wu, Y.; Liu, S.; Chen, Z.; Li, J.; Kanda, N.; Yoshioka, T.; Xiao, X.; et al. 2022 · 2022
Cited alongside, same era.
StyleVC: Non-Parallel Voice Conversion with Adversarial Style Generalization
Hwang, I.-S.; Lee, S.-H.; and Lee, S.-W. 2022 · 2022
Cited alongside, same era.
Flow matching for generative modeling
Lipman, Y.; Chen, R. T.; Ben-Hamu, H.; Nickel, M.; and Le, M. 2022 · 2022
Cited alongside, same era.
DNSMOS P. 835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors
Reddy, C. K.; Gopal, V.; and Cutler, R. 2022 · 2022
Cited alongside, same era.
Utmos: Utokyo-sarulab system for voicemos challenge 2022
Saeki, T.; Xin, D.; Nakata, W.; Koriyama, T.; Takamichi, S.; and Saruwatari, H. 2022 · 2022
Preserving background sound in noise-robust voice conversion via multi-task learning
Yao, J.; Lei, Y.; Wang, Q.; Guo, P.; Ning, Z.; Xie, L.; Li, H.; Liu, J.; and Xie, D. 2023 · 2023
Later among the works it cites.
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
Anastassiou, P.; Tang, Z.; Peng, K.; Jia, D.; Li, J.; Tu, M.; Wang, Y.; Wang, Y.; and Ma, M. 2024 · 2024
Closest in time.
Dddm-vc: Decoupled denoising diffusion models with disentangled representation and prior mixup for verified robust voice conversion
Choi, H.-Y.; Lee, S.-H.; and Lee, S.-W. 2024 · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P.; Kulal, S.; Blattmann, A.; Entezari, R.; Müller, J.; Saini, H.; Levi, Y.; Lorenz, D.; Sauer, A.; Boesel, F.; et al. 2024 · 2024
Closest in time.
Voiceflow: Efficient text-to-speech with rectified flow matching
Guo, Y.; Du, C.; Ma, Z.; Chen, X.; and Yu, K. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Drvc: A framework of any-to-any voice conversion with self-supervised learning
Wang, Q.; Zhang, X.; Wang, J.; Cheng, N.; and Xiao, J. 2022 · 2022
Cited alongside, same era.
Emotional voice conversion: Theory, databases and ESD
Zhou, K.; Sisman, B.; Liu, R.; and Li, H. 2022 · 2022
Cited alongside, same era.
AudioLM: a language modeling approach to audio generation
Borsos, Z.; Marinier, R.; Vincent, D.; Kharitonov, E.; Pietquin, O.; Sharifi, M.; Roblek, D.; Teboul, O.; Grangier, D.; Tagliasacchi, M.; et al. 2023 · 2023
Cited alongside, same era.
Diff-HierVC: Diffusion-based hierarchical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation
Choi, H.-Y.; Lee, S.-H.; and Lee, S.-W. 2023 · 2023
Cited alongside, same era.
ACE-VC: Adaptive and controllable voice conversion using explicitly disentangled self-supervised speech representations
Hussain, S.; Neekhara, P.; Huang, J.; Li, J.; and Ginsburg, B. 2023 · 2023
Cited alongside, same era.
DVQVC: An unsupervised zero-shot voice conversion framework
Li, D.; Li, X.; and Li, X. 2023 · 2023
Cited alongside, same era.
Audioldm: Text-to-audio generation with latent diffusion models
Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; and Plumbley, M. D. 2023 · 2023
Cited alongside, same era.
Closest in time.
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
Jiang, Z.; Liu, J.; Ren, Y.; He, J.; Ye, Z.; Ji, S.; Yang, Q.; Zhang, C.; Wei, P.; Wang, C.; et al. 2024 · 2024
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale
Le, M.; Vyas, A.; Shi, B.; Karrer, B.; Sari, L.; Moritz, R.; Williamson, M.; Manohar, V.; Adi, Y.; Mahadeokar, J.; et al. 2024 · 2024
Closest in time.
DiTTo-TTS: Efficient and Scalable Zero-Shot Text-to-Speech with Diffusion Transformer
Lee, K.; Kim, D. W.; Kim, J.; and Cho, J. 2024 · 2024
Closest in time.
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
Li, J.; Guo, Y.; Chen, X.; and Yu, K. 2024 · 2024
Closest in time.
Posterior Variance-Parameterised Gaussian Dropout: Improving Disentangled Sequential Autoencoders for Zero-Shot Voice Conversion
Luo, Y.-J.; and Dixon, S. 2024 · 2024
Closest in time.
Matcha-TTS: A fast TTS architecture with conditional flow matching
Mehta, S.; Tu, R.; Beskow, J.; Székely, É.; and Henter, G. E. 2024 · 2024
Closest in time.
GR0: Self-Supervised Global Representation Learning for Zero-Shot Voice Conversion
Wang, Y.; Su, J.; Finkelstein, A.; and Jin, Z. 2024 · 2024
Closest in time.
Yang, Y.; Pan, Y.; Yao, J.; Zhang, X.; Ye, J.; Zhou, H.; Xie, L.; Ma, L.; and Zhao, J. 2024 · 2024
Closest in time.
Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts
Yao, J.; Yang, Y.; Lei, Y.; Ning, Z.; Hu, Y.; Pan, Y.; Yin, J.; Zhou, H.; Lu, H.; and Xie, L. 2024 · 2024
Closest in time.