Fetching the paper…
Reading the bibliography…
Any-to-any voice conversion aims to convert the voice from and to any speakers even unseen during training, which is much more challenging compared to one-to-one or many-to-many tasks, but much more attractive in real-world scenarios.
“Unit selection in a concatenative speech synthesis system using a large speech database,”
A. J. Hunt and A. W. Black, · 1996
Earlier work this paper cites.
“Continuous probabilistic transform for voice conversion,”
Y. Stylianou, O. Cappe, and E. Moulines, · 1998
Earlier work this paper cites.
“The ibm trainable speech synthesis system,”
Robert E Donovan and Ellen M Eide, · 1998
Earlier work this paper cites.
“The cmu arctic speech databases,”
John Kominek and Alan W Black, · 2004
Earlier work this paper cites.
“Voice conversion using artificial neural networks,”
S. Desai, E. V. Raghavendra, B. Yegnanarayana, A. W. Black, and K. Prahallad, · 2009
Earlier work this paper cites.
“Exemplar-based voice conversion in noisy environment,”
R. Takashima, T. Takiguchi, and Y. Ariki, · 2012
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“U-net: Convolutional networks for biomedical image segmentation,”
Olaf Ronneberger, Philipp Fischer, and Thomas Brox, · 2015
Earlier work this paper cites.
“Cute: A concatenative method for voice conversion using exemplar-based unit selection,”
Z. Jin, A. Finkelstein, S. DiVerdi, J. Lu, and G. J. Mysore, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” 2017
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al., · 2017
Cited alongside, same era.
“Multi-target voice conversion without parallel data by adversarially learning disentangled audio representations,”
Ju chieh Chou, Cheng chieh Yeh, Hung yi Lee, and Lin shan Lee, · 2018
Cited alongside, same era.
“Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks,”
T. Kaneko and H. Kameoka, · 2018
Cited alongside, same era.
“Stargan-vc: Non-parallel many-to-many voice conversion using star generative adversarial networks,”
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, and Nobukatsu Hojo, · 2018
Cited alongside, same era.
“Voice transformer network: Sequence-to-sequence voice conversion using transformer with text-to-speech pretraining,” 2019
Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu, Hirokazu Kameoka, and Tomoki Toda, · 2019
Later among the works it cites.
“One-Shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization,”
Ju chieh Chou and Hung-Yi Lee, · 2019
Later among the works it cites.
“Autovc: Zero-shot voice style transfer with only autoencoder loss,”
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson, · 2019
Later among the works it cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2019
Later among the works it cites.
“Towards Achieving Robust Universal Neural Vocoding,”
Jaime Lorenzo-Trueba, Thomas Drugman, Javier Latorre, Thomas Merritt, Bartosz Putrycz, Roberto Barra-Chicote, Alexis Moinet, and Vatsal Aggarwal, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, · 2018
Cited alongside, same era.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, and Yonghui Wu, · 2018
Cited alongside, same era.
“Generalized end-to-end loss for speaker verification,”
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno, · 2018
Cited alongside, same era.
“Atts2s-vc: Sequence-to-sequence voice conversion with attention and context preservation mechanisms,”
K. Tanaka, H. Kameoka, T. Kaneko, and N. Hojo, · 2019
Cited alongside, same era.
“Libritts: A corpus derived from librispeech for text-to-speech,” 2019
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Later among the works it cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,” 2020
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Closest in time.
“Voice conversion with transformer network,”
R. Liu, X. Chen, and X. Wen, · 2020
Closest in time.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,” 2020
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2020
Closest in time.