Fetching the paper…
Reading the bibliography…
Non-parallel many-to-many voice conversion, as well as zero-shot voice conversion, remain under-explored areas.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Kingma, D. P., Mohamed, S., Rezende, D. J., and Welling, M · 2014
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
End-to-end text-dependent speaker verification
Heigold, G., Moreno, I., Bengio, S., and Shazeer, N · 2016
Earlier work this paper cites.
Voice conversion from non-parallel corpora using variational auto-encoder
Hsu, C.-C., Hwang, H.-T., Wu, Y.-C., Tsao, Y., and Wang, H.-M · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Van Den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Superseded-CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit
Veaux, C., Yamagishi, J., MacDonald, K., et al · 2016
Earlier work this paper cites.
Analysis of the voice conversion challenge 2016 evaluation results
Wester, M., Wu, Z., and Yamagishi, J · 2016
Earlier work this paper cites.
A KL divergence and DNN-based approach to voice conversion without parallel training sentences
Xie, F.-L., Soong, F. K., and Li, H · 2016
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech
Arik, S. O., Chrzanowski, M., Coates, A., Diamos, G., Gibiansky, A., Kang, Y., Li, X., Miller, J., Ng, A., Raiman, J., et al · 2017
Earlier work this paper cites.
Hsu, C.-C., Hwang, H.-T., Wu, Y.-C., Tsao, Y., and Wang, H.-M · 2017
Cited alongside, same era.
Parallel-data-free voice conversion using cycle-consistent adversarial networks
Kaneko, T. and Kameoka, H · 2017
Cited alongside, same era.
Voxceleb: a large-scale speaker identification dataset
Nagrani, A., Chung, J. S., and Zisserman, A · 2017
Cited alongside, same era.
Segan: Speech enhancement generative adversarial network
Pascual, S., Bonafonte, A., and Serra, J · 2017
Cited alongside, same era.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Zhu, J.-Y., Park, T., Isola, P., and Efros, A. A · 2017
Cited alongside, same era.
A multi-discriminator cyclegan for unsupervised non-parallel speech domain adaptation
Hosseini-Asl, E., Zhou, Y., Xiong, C., and Socher, R · 2018
Later among the works it cites.
Voice conversion based on cross-domain features using variational auto encoders
Huang, W.-C., Hwang, H.-T., Peng, Y.-H., Tsao, Y., and Wang, H.-M · 2018
Later among the works it cites.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Jia, Y., Zhang, Y., Weiss, R. J., Wang, Q., Shen, J., Ren, F., Chen, Z., Nguyen, P., Pang, R., Moreno, I. L., et al · 2018
Later among the works it cites.
Non-parallel voice conversion using variational autoencoders conditioned by phonetic posteriorgrams and d-vectors
Saito, Y., Ijima, Y., Nishida, K., and Takamichi, S · 2018
Later among the works it cites.
Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chou, J.-c., Yeh, C.-c., Lee, H.-y., and Lee, L.-s · 2018
Cited alongside, same era.
Donahue, C., McAuley, J., and Puckette, M · 2018
Cited alongside, same era.
SVSGAN: Singing voice separation via generative adversarial network
Fan, Z.-C., Lai, Y.-L., and Jang, J.-S. R · 2018
Cited alongside, same era.
High-quality nonparallel voice conversion based on cycle-consistent adversarial network
Fang, F., Yamagishi, J., Echizen, I., and Lorenzo-Trueba, J · 2018
Cited alongside, same era.
Voice impersonation using generative adversarial networks
Gao, Y., Singh, R., and Raj, B · 2018
Cited alongside, same era.
Kameoka, H., Kaneko, T., Tanaka, K., and Hojo, N
Cited in the paper.
Stargan-vc: Non-parallel many-to-many voice conversion with star generative adversarial networks
Kameoka, H., Kaneko, T., Tanaka, K., and Hojo, N
Cited in the paper.
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerrv-Ryan, R., et al · 2018
Later among the works it cites.
Generative adversarial source separation
Subakan, Y. C. and Smaragdis, P · 2018
Later among the works it cites.
Generalized end-to-end loss for speaker verification
Wan, L., Wang, Q., Papir, A., and Moreno, I. L · 2018
Later among the works it cites.
Look ma, no GANs! image transformation with modifAE, 2019
Atalla, C., Tam, B., Song, A., and Cottrell, G · 2019
Closest in time.
Biadsy, F., Weiss, R. J., Moreno, P. J., Kanvesky, D., and Jia, Y · 2019
Closest in time.
Unsupervised singing voice conversion
Nachmani, E. and Wolf, L · 2019
Closest in time.