Fetching the paper…
Reading the bibliography…
One-shot voice conversion (VC) with only a single target speaker's speech for reference has become a hot research topic.
A. Packman, M. Onslow, and R. Menzies, “Novel speech patterns and the treatment of stuttering,” Disability and Rehabilitation , vol. 22, no. 1-2, pp. 65–79, 2000
2000
Earlier work this paper cites.
D. Gibbon and U. Gut, “Measuring speech rhythm,” in Seventh European Conference on Speech Communication and Technology , 2001
2001
Earlier work this paper cites.
E. E. Helander and J. Nurminen, “A novel method for prosody prediction in voice conversion,” in 2007 IEEE International Conference on Acoustics, Speech and Signal Processing-ICASSP’07 , vol. 4. IEEE, 2007, pp. IV–509
2007
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 8, pp. 2222–2235, 2007
2007
Earlier work this paper cites.
Nahler and Gerhard, “Pearson correlation coefficient,” Springer Vienna , vol. 10.1007/978-3-211-89836-9, no. Chapter 1025, pp. 132–132, 2009
2009
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky, “Domain-adversarial training of neural networks,” Journal of Machine Learning Research , vol. 17, no. 1, pp. 2096–2030, 2016
2016
Earlier work this paper cites.
M. K. Veaux Christophe, Yamagishi Junichi, “Superseded - cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” 2016
2016
Earlier work this paper cites.
A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio.” SSW , vol. 125, p. 2, 2016
2016
Earlier work this paper cites.
S. H. Mohammadi and A. Kain, “An overview of voice conversion systems,” Speech Communication , vol. 88, pp. 65–82, 2017
2017
Earlier work this paper cites.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice Conversion from Unaligned Corpora Using Variational Autoencoding Wasserstein Generative Adversarial Networks,” in Proc. Interspeech 2017 , 2017, pp. 3364–3368
2017
Earlier work this paper cites.
S. Liu, J. Zhong, L. Sun, X. Wu, X. Liu, and H. Meng, “Voice Conversion Across Arbitrary Speakers Based on a Single Target-Speaker Utterance,” in Proc. Interspeech 2018 , 2018, pp. 496–500
2018
Cited alongside, same era.
F. Fang, J. Yamagishi, I. Echizen, and J. Lorenzo-Trueba, “High-quality nonparallel voice conversion based on cycle-consistent adversarial network,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5279–5283
2018
Cited alongside, same era.
H. Lu, Z. Wu, D. Dai, R. Li, S. Kang, J. Jia, and H. Meng, “One-Shot Voice Conversion with Global Speaker Embeddings,” in Proc. Interspeech 2019 , 2019, pp. 669–673
2019
Cited alongside, same era.
J. Chou and H. Lee, “One-shot voice conversion by separating speaker and content representations with instance normalization,” in Interspeech , 2019, pp. 664–668
2019
Cited alongside, same era.
P. Cheng, W. Hao, S. Dai, J. Liu, Z. Gan, and L. Carin, “Club: A contrastive log-ratio upper bound of mutual information,” in Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event , ser. Proceedings of Machine Learning Research, vol. 119. PMLR, 2020, pp. 1779–1788
2020
Later among the works it cites.
K. Qian, Y. Zhang, S. Chang, M. Hasegawa-Johnson, and D. Cox, “Unsupervised speech decomposition via triple information bottleneck,” in Proceedings of the 37th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, H. D. III and A. Singh, Eds., vol. 119. PMLR, 13–18 Jul 2020, pp. 7836–7846
2020
Later among the works it cites.
B. Sisman, J. Yamagishi, S. King, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 132–157, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “Autovc: Zero-shot voice style transfer with only autoencoder loss,” in International Conference on Machine Learning . PMLR, 2019, pp. 5210–5219
2019
Cited alongside, same era.
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “Acvae-vc: Non-parallel voice conversion with auxiliary classifier variational autoencoder,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 9, pp. 1432–1443, 2019
2019
Cited alongside, same era.
A. T. Liu, P. chun Hsu, and H.-Y. Lee, “Unsupervised End-to-End Learning of Discrete Linguistic Units for Voice Conversion,” in Proc. Interspeech 2019 , 2019, pp. 1108–1112
2019
Cited alongside, same era.
A. Polyak and L. Wolf, “Attention-based wavenet autoencoder for universal voice conversion,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6800–6804
2019
Cited alongside, same era.
K. Qian, Z. Jin, M. Hasegawa-Johnson, and G. J. Mysore, “F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6284–6288
2020
Cited alongside, same era.
D. Wu, Y. Chen, and H. Lee, “Vqvc+: One-shot voice conversion by vector quantization and u-net architecture,” in Interspeech , 2020, p. 4691–4695
2020
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, M. Hasegawa-Johnson, and D. Cox, “Unsupervised speech decomposition via triple information bottleneck,” in International Conference on Machine Learning . PMLR, 2020, pp. 7836–7846
2020
Cited alongside, same era.
2021
Later among the works it cites.
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “CycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-Spectrogram Conversion,” in Proc. Interspeech 2020 , 2020, pp. 2017–2021
2021
Later among the works it cites.
Z. Lian, R. Zhong, Z. Wen, B. Liu, and J. Tao, “Towards fine-grained prosody control for voice conversion,” in 2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP) , 2021, pp. 1–5
2021
Later among the works it cites.
J. Wang, J. Li, X. Zhao, Z. Wu, S. Kang, and H. Meng, “Adversarially Learning Disentangled Speech Representations for Robust Multi-Factor Voice Conversion,” in Proc. Interspeech 2021 , 2021, pp. 846–850
2021
Later among the works it cites.
D. Wang, L. Deng, Y. T. Yeung, X. Chen, X. Liu, and H. Meng, “VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-Shot Voice Conversion,” in Interspeech , 2021, pp. 1344–1348
2021
Later among the works it cites.
H. Tang, X. Zhang, J. Wang, N. Cheng, and J. Xiao, “Clsvc: Learning speech representations with two different classification tasks.” Openreview, 2021, https://openreview.net/forum?id=xp2D-1PtLc5
2021
Later among the works it cites.
2021
Later among the works it cites.
C. Ho Chan, K. Qian, Y. Zhang, and M. Hasegawa-Johnson, “Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottlenecks,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 6332–6336
2022
Closest in time.