Fetching the paper…
Reading the bibliography…
Recent developments in neural speech synthesis and vocoding have sparked a renewed interest in voice conversion (VC).
R. Carlson, A. Friberg, L. Fryden, B. Granstrm, and J. Sundberg, “Speech and music performance: Parallels and contrasts,” Contemporary Music Review , vol. 4, pp. 391–404, 1989
1989
Earlier work this paper cites.
K. Kasi and S. A. Zahorian, “Yet another algorithm for pitch tracking,” in Proc. ICASSP , Orlando, FL, USA, 2002, pp. I–361–I–364
2002
Earlier work this paper cites.
U. Zölzer, X. Amatriain, D. Arfib, J. Bonada, G. De Poli, P. Dutilleux, G. Evangelista, F. Keiler, A. Loscos, D. Rocchesso et al. , DAFX: Digital audio effects . John Wiley & Sons, 2002
2002
Earlier work this paper cites.
Y.-Y. Wang, A. Acero, and C. Chelba, “Is word error rate a good indicator for spoken language understanding accuracy,” in Proc. ASRU , Virgin Islands, 2003, pp. 577–582
2003
Earlier work this paper cites.
L. Chen, Z. Ling, W. Guo, and L. Dai, “GMM-based voice conversion with explicit modelling on feature transform,” in Proc. ISCSLP , Tainan, Taiwan, 2010, pp. 364–368
2010
Earlier work this paper cites.
B. Schuller and A. Batliner, Computational paralinguistics: emotion, affect and personality in speech and language processing . John Wiley & Sons, 2013
2013
Earlier work this paper cites.
L. Sun, S. Kang, K. Li, and H. Meng, “Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,” in Proc. ICASSP , South Brisbane, Australia, 2015, pp. 4869–4873
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in Proc. ICASSP , South Brisbane, Australia, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. C. Streijl, S. Winkler, and D. S. Hands, “Mean opinion score (MOS) revisited: methods and applications, limitations and alternatives,” Multimedia Systems , vol. 22, no. 2, pp. 213–227, 2016
2016
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “VoxCeleb: a large-scale speaker identification dataset,” in Proc. INTERSPEECH , Stockholm, Sweden, 2017
2017
Earlier work this paper cites.
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “StarGAN-VC: Non-parallel many-to-many voice conversion using star generative adversarial networks,” in Proc. SLT , Athens, Greece, 2018, pp. 266–273
2018
Cited alongside, same era.
F. Charpentier and M. Stella, “Diphone synthesis using an overlap-add technique for speech waveforms concatenation,” in Proc. ICASSP , Tokyo, Japan, 1986, pp. 2015–2018
2018
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-shot voice style transfer with only Autoencoder loss,” in Proc. ICML , Long Beach, USA, 2019, pp. 5210–5219
2019
Cited alongside, same era.
J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English multi-speaker corpus for CSTR Voice Cloning Toolkit (version 0.92),” https://doi.org/10.7488/ds/2645 , University of Edinburgh, The Centre for Speech Technology Research (CSTR), 2019
2019
Cited alongside, same era.
Y. Y. Lin, C. Chien, J. Lin, H. Lee, and L. Lee, “FragmentVC: Any-to-any voice conversion by end-to-end extracting and fusing fine-grained voice fragments with attention,” in Proc. ICASSP , Toronto, Canada, 2021, pp. 5939–5943
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Lian, R. Zhong, Z. Wen, B. Liu, and J. Tao, “Towards fine-grained prosody control for voice conversion,” in Proc. ISCSLP , Hong Kong, 2021, pp. 1–5
2021
Later among the works it cites.
H.-S. Choi, J. Lee, W. Kim, J. Lee, H. Heo, and K. Lee, “Neural analysis and synthesis: Reconstructing speech from self-supervised representations,” Advances in Neural Information Processing Systems , vol. 34, pp. 16 251–16 265, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Sisman, J. Yamagishi, S. King, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,” IEEE Trans. ASLP , vol. 29, pp. 132–157, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, M. Hasegawa-Johnson, and D. Cox, “Unsupervised speech decomposition via triple information bottleneck,” in Proc. ICML , Virtual, 2020, pp. 7836–7846
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Y. A. Li, A. Zare, and N. Mesgarani, “StarGANv2-VC: A diverse, unsupervised, non-parallel framework for natural-sounding voice conversion,” in Proc. INTERSPEECH , Brno, Czechia, 2021
2021
Cited alongside, same era.
S. Nercessian, “End-to-end zero-shot voice conversion using a DDSP vocoder,” in Proc. WASPAA , New Paltz, NY, USA, 2021, pp. 1–5
2021
Cited alongside, same era.
Meta Research, “Facebook AI Research Sequence-to-Sequence Toolkit written in Python,” https://github.com/facebookresearch/fairseq , GitHub
Cited in the paper.
K. Qian, Z. Jin, M. Hasegawa-Johnson, and G. J. Mysore, “F0-consistent many-to-many non-parallel voice conversion via conditional Autoencoder,” in Proc. INTERSPEECH , Incheon, Korea, 2022, pp. 6284–6288
2022
Closest in time.
S. Wang, D. Kostadinov, and D. Borth, “Zero-shot voice conversion via self-supervised prosody representation learning,” in Proc. IJCNN , Padua, Italy, 2022, pp. 01–08
2022
Closest in time.
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,” in Proc. ICML , Hyderabad, India, 2022, pp. 2709–2720
2022
Closest in time.
B. Nguyen and F. Cardinaux, “NVC-Net: End-to-end adversarial voice conversion,” in Proc. ICASSP , Singapore, 2022, pp. 7012–7016
2022
Closest in time.
L. Chen and A. Rudnicky, “Fine-grained style control in transformer-based text-to-speech synthesis,” in Proc. ICASSP , Singapore, 2022, pp. 7907–7911
2022
Closest in time.
S. Lee, H. Noh, W. Nam, and S. Lee, “Duration controllable voice conversion via phoneme-based information bottleneck,” IEEE Trans. ASLP , vol. 30, pp. 1173–1183, 2022
2022
Closest in time.