Fetching the paper…
Reading the bibliography…
Voice conversion (VC) can be achieved by first extracting source content information and target speaker information, and then reconstructing waveform with these information.
“Pearson correlation coefficient,”
J Benesty, J Chen, Y Huang, et al., · 2009
Earlier work this paper cites.
“Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,”
L Sun, K Li, H Wang, et al., · 2016
Earlier work this paper cites.
“Density estimation using real nvp,”
L Dinh, J Sohl-Dickstein, and S Bengio, · 2016
Earlier work this paper cites.
“Autoencoding beyond pixels using a learned similarity metric,”
A. B. L Larsen, S. K Sønderby, H Larochelle, et al., · 2016
Earlier work this paper cites.
“An overview of voice conversion systems,”
S. H Mohammadi and A Kain, · 2017
Earlier work this paper cites.
“Least squares generative adversarial networks,”
X Mao, Q Li, H Xie, et al., · 2017
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Y Wang, D Stanton, Y Zhang, et al., · 2018
Earlier work this paper cites.
“Collapsed speech segment detection and suppression for wavenet vocoder,”
Y.-C Wu, K Kobayashi, T Hayashi, et al., · 2018
Earlier work this paper cites.
“Autovc: Zero-shot voice style transfer with only autoencoder loss,”
K Qian, Y Zhang, S Chang, et al., · 2019
Earlier work this paper cites.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),”
J Yamagishi, C Veaux, K MacDonald, et al., · 2019
Earlier work this paper cites.
“Libritts: A corpus derived from librispeech for text-to-speech,”
H Zen, V Dang, R Clark, et al., · 2019
Cited alongside, same era.
“Cotatron: Transcription-guided speech encoder for any-to-many voice conversion without parallel data,”
S.-w Park, D.-y Kim, and M.-c Joe, · 2020
Cited alongside, same era.
“Vqvc+: One-shot voice conversion by vector quantization and u-net architecture,”
D.-Y Wu, Y.-H Chen, and H.-y Lee, · 2020
Cited alongside, same era.
“Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion,”
Y Zhao, W.-C Huang, X Tian, et al., · 2020
Cited alongside, same era.
“Unsupervised cross-lingual representation learning for speech recognition,”
A Conneau, A Baevski, R Collobert, et al., · 2020
“Again-vc: A one-shot voice conversion using activation guidance and adaptive instance normalization,”
Y.-H Chen, D.-Y Wu, T.-H Wu, et al., · 2021
Later among the works it cites.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
J Kim, J Kong, and J Son, · 2021
Later among the works it cites.
“Neural analysis and synthesis: Reconstructing speech from self-supervised representations,”
H.-S Choi, J Lee, W Kim, et al., · 2021
Later among the works it cites.
D Wang, L Deng, Y. T Yeung, et al., · 2021
Later among the works it cites.
“Large-scale self-supervised speech representation learning for automatic speaker verification,”
Z Chen, S Chen, Y Wu, et al., · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
J Kong, J Kim, et al., · 2020
Cited alongside, same era.
“Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,”
K Zhou, B Sisman, R Liu, et al., · 2021
Cited alongside, same era.
“Any-to-many voice conversion with location-relative sequence-to-sequence modeling,”
S Liu, Y Cao, D Wang, et al., · 2021
Cited alongside, same era.
“Transfer learning from speech synthesis to voice conversion with non-parallel training data,”
M Zhang, Y Zhou, L Zhao, et al., · 2021
Cited alongside, same era.
Closest in time.
“S3prl-vc: Open-source voice conversion framework with self-supervised speech representations,”
W.-C Huang, S.-W Yang, T Hayashi, et al., · 2022
Closest in time.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
S Chen, C Wang, Z Chen, et al., · 2022
Closest in time.
“Speechsplit2. 0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottlenecks,”
C. H Chan, K Qian, Y Zhang, et al., · 2022
Closest in time.
“Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,”
E Casanova, J Weber, C. D Shulby, et al., · 2022
Closest in time.