Fetching the paper…
Reading the bibliography…
Any-to-any voice conversion aims to transform source speech into a target voice with just a few examples of the target speaker as a reference.
E. Fix, Discriminatory analysis. Nonparametric discrimination: Consistency properties . USAF school of Aviation Medicine, 1985, vol. 1
1985
Earlier work this paper cites.
D. Sundermann, H. Hoge, A. Bonafonte, H. Ney, A. Black, and S. Narayanan, “Text-independent voice conversion based on unit selection,” in ICASSP , 2006
2006
Earlier work this paper cites.
K. Fujii, J. Okawa, and K. Suigetsu, “High-individuality voice conversion based on concatenative speech synthesis,” International Journal of Electrical and Computer Engineering , vol. 1, no. 11, pp. 1625 -- 1630, 2007
2007
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: An ASR corpus based on public domain audio books,” in ICASSP , 2015
2015
Earlier work this paper cites.
Z. Jin, A. Finkelstein, S. DiVerdi, J. Lu, and G. J. Mysore, “Cute: A concatenative method for voice conversion using exemplar-based unit selection,” in ICASSP , 2016
2016
Earlier work this paper cites.
S. H. Mohammadi and A. Kain, “An overview of voice conversion systems,” Speech Communication , vol. 88, pp. 65–82, 2017
2017
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-Vectors: Robust dnn embeddings for speaker recognition,” in ICASSP , 2018
2018
Earlier work this paper cites.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-shot voice style transfer with only autoencoder loss,” in PMLR , 2019
2019
Earlier work this paper cites.
J.-c. Chou and H.-Y. Lee, “One-shot voice conversion by separating speaker and content representations with instance normalization,” in Interspeech , 2019
2019
Cited alongside, same era.
Z. Yi, W.-C. Huang, X. Tian, J. Yamagishi, R. K. Das, T. Kinnunen, Z.-H. Ling, and T. Toda, “Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion,” in Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge , 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NeurIPS , 2020
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in NeurIPS , 2020
2020
Cited alongside, same era.
2022
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Later among the works it cites.
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,” in PMLR , 2022
2022
Later among the works it cites.
E. Dunbar, N. Hamilakis, and E. Dupoux, “Self-supervised language learning from raw audio: Lessons from the zero resource speech challenge,” Journal of Selected Topics in Signal Processing , 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Liu, Y. Cao, D. Wang, X. Wu, X. Liu, and H. Meng, “Any-to-many voice conversion with location-relative sequence-to-sequence modeling,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1717–1728, 2021
2021
Cited alongside, same era.
D. Wang, L. Deng, Y. T. Yeung, X. Chen, X. Liu, and H. Meng, “VQMIVC: Vector quantization and mutual information-based unsupervised speech representation disentanglement for one-shot voice conversion,” in Interspeech , 2021
2021
Cited alongside, same era.
A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” in IEEE ASRU , 2021
2021
Cited alongside, same era.
B. Sisman, J. Yamagishi, S. King, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 132–157, 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
B. van Niekerk, M.-A. Carbonneau, J. Zaïdi, M. Baas, H. Seuté, and H. Kamper, “A comparison of discrete and soft speech units for improved voice conversion,” in ICASSP , 2022
2022
Later among the works it cites.
G.-T. Lin, C.-L. Feng, W.-P. Huang, Y. Tseng, T.-H. Lin, C.-A. Li, H.-y. Lee, and N. G. Ward, “On the utility of self-supervised models for prosody-related tasks,” in IEEE SLT , 2023
2023
Closest in time.