Fetching the paper…
Reading the bibliography…
Voice conversion (VC) is a task that transforms voice from target audio to source without losing linguistic contents, it is challenging especially when source and target speakers are unseen during training (zero-shot VC).
O. Türk and L. Arslan, “Robust processing techniques for voice conversion,” Comput. Speech Lang. , vol. 20, pp. 441–467, 2006
2006
Earlier work this paper cites.
W. J. Hess, “Pitch and voicing determination of speech with an extension toward music signals,” Springer handbook of speech processing , pp. 181–212, 2008
2008
Earlier work this paper cites.
L. V. D. Maaten and G. E. Hinton, “Visualizing data using t-sne,” Journal of Machine Learning Research , vol. 9, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
E. Helander, T. Virtanen, J. Nurminen, and M. Gabbouj, “Voice conversion using partial least squares regression,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, pp. 912–921, 2010
2010
Earlier work this paper cites.
Y. Qian, J. Xu, and F. Soong, “A frame mapping based hmm approach to cross-lingual voice transformation,” 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 5120–5123, 2011
2011
Earlier work this paper cites.
L. Chen, Z. Ling, L. Liu, and L.-R. Dai, “Voice conversion using deep neural networks with layer-wise generative training,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 22, pp. 1859–1872, 2014
2014
Earlier work this paper cites.
L. Sun, S. Kang, K. Li, and H. Meng, “Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 4869–4873, 2015
2015
Earlier work this paper cites.
L. Sukhostat and Y. Imamverdiyev, “A comparative analysis of pitch detection methods under the influence of different noise conditions,” Journal of voice , vol. 29, no. 4, pp. 410–417, 2015
2015
Earlier work this paper cites.
L. Sun, H. Wang, S. Kang, K. Li, and H. Meng, “Personalized, cross-lingual tts using phonetic posteriorgrams,” in INTERSPEECH , 2016
2016
Earlier work this paper cites.
F. Xie, F. Soong, and H. Li, “A kl divergence and dnn-based approach to voice conversion without parallel training sentences,” in INTERSPEECH , 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
D. Ulyanov, A. Vedaldi, and V. Lempitsky, “Improved texture networks: Maximizing quality and diversity in feed-forward stylization and texture synthesis,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 4105–4113, 2017
2017
Cited alongside, same era.
A. Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in NIPS , 2017
2017
Cited alongside, same era.
C. Veaux, J. Yamagishi, and K. Macdonald, “Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” 2017
2017
Cited alongside, same era.
——, “Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks,” 2018 26th European Signal Processing Conference (EUSIPCO) , pp. 2100–2104, 2018
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using wavenet autoencoders,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, pp. 2041–2053, 2019
2019
Later among the works it cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” in INTERSPEECH , 2019
2019
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “Stargan-vc: non-parallel many-to-many voice conversion using star generative adversarial networks,” 2018 IEEE Spoken Language Technology Workshop (SLT) , pp. 266–273, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Serrà, S. Pascual, and C. Segura, “Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion,” in NeurIPS , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J.-C. Chou, C. chieh Yeh, and H. yi Lee, “One-shot voice conversion by separating speaker and content representations with instance normalization,” in INTERSPEECH , 2019
2019
Cited alongside, same era.
M. Zhang, X. Wang, F. Fang, H. Li, and J. Yamagishi, “Joint training framework for text-to-speech and voice conversion using multi-source tacotron and wavenet,” in INTERSPEECH , 2019
2019
Cited alongside, same era.
D. Wu and H. yi Lee, “One-shot voice conversion by vector quantization,” ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 7734–7738, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
B. Niekerk, L. Nortje, and H. Kamper, “Vector-quantized neural networks for acoustic unit discovery in the zerospeech 2020 challenge,” in INTERSPEECH , 2020
2020
Later among the works it cites.
R. Yamamoto, E. Song, and J. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6199–6203, 2020
2020
Later among the works it cites.