Fetching the paper…
Reading the bibliography…
We present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes.
“Acoustic theory of speech production,”
G. Fant, · 1960
Earlier work this paper cites.
“Prosody in the comprehension of spoken language: A literature review,”
A. Cutler, D. Dahan, and W. van Donselaar, · 1997
Earlier work this paper cites.
“Visualizing data using t-sne,”
L. van der Maaten and G. Hinton, · 2008
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, · 2016
Earlier work this paper cites.
“Least squares generative adversarial networks,”
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley, · 2017
Earlier work this paper cites.
“Crepe: A convolutional representation for pitch estimation,”
J. W. Kim, J. Salamon, P. Li, and J. P. Bello, · 2018
Earlier work this paper cites.
“AutoVC: Zero-shot voice style transfer with only autoencoder loss,”
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, · 2019
Earlier work this paper cites.
“Unsupervised End-to-End Learning of Discrete Linguistic Units for Voice Conversion,”
A. T. Liu, P. chun Hsu, and H.-Y. Lee, · 2019
Earlier work this paper cites.
“wav2vec: Unsupervised Pre-Training for Speech Recognition,”
S. Schneider, A. Baevski, R. Collobert, and M. Auli, · 2019
Earlier work this paper cites.
“Fastspeech: Fast, robust and controllable text to speech,”
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2019
Earlier work this paper cites.
“CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),” 2019
J. Yamagishi, C. Veaux, and K. MacDonald, · 2019
Cited alongside, same era.
“LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,”
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, · 2019
Cited alongside, same era.
“F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder,”
K. Qian, Z. Jin, M. Hasegawa-Johnson, and G. J. Mysore, · 2020
Cited alongside, same era.
“VQVC+: One-Shot Voice Conversion by Vector Quantization and U-Net Architecture,”
D.-Y. Wu, Y.-H. Chen, and H. yi Lee, · 2020
Cited alongside, same era.
“Unsupervised speech decomposition via triple information bottleneck,”
K. Qian, Y. Zhang, S. Chang, M. Hasegawa-Johnson, and D. Cox, · 2020
Cited alongside, same era.
“Neural analysis and synthesis: Reconstructing speech from self-supervised representations,”
H.-S. Choi, J. Lee, W. Kim, J. Lee, H. Heo, and K. Lee, · 2021
Later among the works it cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, · 2021
Later among the works it cites.
“Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,”
K. Zhou, B. Sisman, R. Liu, and H. Li, · 2021
Later among the works it cites.
“SpeechSplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottlenecks,”
C. Ho Chan, K. Qian, Y. Zhang, and M. Hasegawa-Johnson, · 2022
Closest in time.
“Speech Representation Disentanglement with Adversarial Mutual Information Learning for One-shot Voice Conversion,”
S. Yang, M. Tantrawenith, H. Zhuang, Z. Wu, A. Sun, J. Wang, N. Cheng, H. Tang, X. Zhao, J. Wang, and H. Meng, · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, · 2020
Cited alongside, same era.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
J. Kong, J. Kim, and J. Bae, · 2020
Cited alongside, same era.
“Espnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,”
T. Hayashi, R. Yamamoto, K. Inoue, T. Yoshimura, S. Watanabe, T. Toda, K. Takeda, Y. Zhang, and X. Tan, · 2020
Cited alongside, same era.
“VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-Shot Voice Conversion,”
D. Wang, L. Deng, Y. T. Yeung, X. Chen, X. Liu, and H. Meng, · 2021
Cited alongside, same era.
“Global prosody style transfer without text transcriptions,”
K. Qian, Y. Zhang, S. Chang, J. Xiong, C. Gan, D. Cox, and M. Hasegawa-Johnson, · 2021
Cited alongside, same era.
Closest in time.
“S3PRL-VC: Open-source voice conversion framework with self-supervised speech representations,”
W.-C. Huang, S.-W. Yang, T. Hayashi, H.-Y. Lee, S. Watanabe, and T. Toda, · 2022
Closest in time.
“Fine-grained style control in transformer-based text-to-speech synthesis,”
L.-W. Chen and A. Rudnicky, · 2022
Closest in time.
“textless-lib: a library for textless spoken language processing,”
E. Kharitonov, J. Copet, K. Lakhotia, T. A. Nguyen, P. Tomasello, A. Lee, A. Elkahky, W.-N. Hsu, A. Mohamed, E. Dupoux, and Y. Adi, · 2022
Closest in time.
“Text-free prosody-aware generative spoken language modeling,”
E. Kharitonov, A. Lee, A. Polyak, Y. Adi, J. Copet, K. Lakhotia, T. A. Nguyen, M. Riviere, A. Mohamed, E. Dupoux, and W.-N. Hsu, · 2022
Closest in time.
“Dawn of the transformer era in speech emotion recognition: closing the valence gap,”
J. Wagner, A. Triantafyllopoulos, H. Wierstorf, M. Schmitt, F. Burkhardt, F. Eyben, and B. W. Schuller, · 2022
Closest in time.