Fetching the paper…
Reading the bibliography…
This paper proposes a method that allows non-parallel many-to-many voice conversion (VC) by using a variant of a generative adversarial network (GAN) called StarGAN.
“Spectral voice conversion for text-to-speech synthesis,”
A. Kain and M. W. Macon, · 1998
Earlier work this paper cites.
“Continuous probabilistic transform for voice conversion,”
Y. Stylianou, O. Cappé, and E. Moulines, · 1998
Earlier work this paper cites.
“Improving the intelligibility of dysarthric speech,”
A. B. Kain, J.-P. Hosom, X. Niu, J. P. van Santen, M. Fried-Oken, and J. Staehely, · 2007
Earlier work this paper cites.
“Voice conversion based on maximumlikelihood estimation of spectral parameter trajectory,”
T. Toda, A. W. Black, and K. Tokuda, · 2007
Earlier work this paper cites.
“High quality voice conversion through phoneme-based linear mapping functions with STRAIGHT for mandarin,”
K. Liu, J. Zhang, and Y. Yan, · 2007
Earlier work this paper cites.
“Data-driven emotion conversion in spoken English,”
Z. Inanoglu and S. Young, · 2009
Earlier work this paper cites.
“Evaluation of expressive speech synthesis with voice conversion and copy resynthesis techniques,”
O. Türk and M. Schröder, · 2010
Earlier work this paper cites.
“Voice conversion using partial least squares regression,”
E. Helander, T. Virtanen, J. Nurminen, and M. Gabbouj, · 2010
Earlier work this paper cites.
“Spectral mapping using artificial neural networks for voice conversion,”
S. Desai, A. W. Black, B. Yegnanarayana, and K. Prahallad, · 2010
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, · 2011
Earlier work this paper cites.
“Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speech,”
K. Nakamura, T. Toda, H. Saruwatari, and K. Shikano, · 2012
Earlier work this paper cites.
“Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,”
T. Toda, M. Nakagiri, and K. Shikano, · 2012
Earlier work this paper cites.
“Exemplar-based voice conversion using sparse representation in noisy environments,”
R. Takashima, T. Takiguchi, and Y. Ariki, · 2013
Earlier work this paper cites.
“Voice conversion using deep neural networks with speaker-independent pre-training,”
S. H. Mohammadi and A. Kain, · 2014
Earlier work this paper cites.
“Exemplar-based sparse representation with residual compensation for voice conversion,”
Z. Wu, T. Virtanen, E. S. Chng, and H. Li, · 2014
Earlier work this paper cites.
“Voice conversion using deep neural networks with layer-wise generative training,”
L.-H. Chen, Z.-H. Ling, L.-J. Liu, and L.-R. Dai, · 2014
Earlier work this paper cites.
“Voice conversion based on speaker-dependent restricted Boltzmann machines,”
T. Nakashika, T. Takiguchi, and Y. Ariki, · 2014
Earlier work this paper cites.
“High-order sequence modeling using speaker-dependent recurrent temporal restricted Boltzmann machines for voice conversion,”
T. Nakashika, T. Takiguchi, and Y. Ariki, · 2014
Cited alongside, same era.
“Auto-encoding variational Bayes,”
D. P. Kingma and M. Welling, · 2014
Cited alongside, same era.
“Semi-supervised learning with deep generative models,”
D. P. Kingma and D. J. Rezendey, S. Mohamedy, and M. Welling, · 2014
Cited alongside, same era.
“Generative adversarial nets,”
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, · 2014
Cited alongside, same era.
“Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,”
L. Sun, S. Kang, K. Li, and H. Meng, · 2015
Cited alongside, same era.
“Parallel-data-free, many-to-many voice conversion using an adaptive restricted Boltzmann machine,”
“Neural discrete representation learning,”
A. van den Oord and O. Vinyals, · 2017
Later among the works it cites.
“Parallel-data-free many-to-many voice conversion based on dnn integrated with eigenspace using a non-parallel speech corpus,”
T. Hashimoto, H. Uchida, D. Saito, and N. Minematsu, · 2017
Later among the works it cites.
“Generative adversarial network-based postfilter for statistical parametric speech synthesis,”
T. Kaneko, H. Kameoka, N. Hojo, Y. Ijima, K. Hiramatsu, and K. Kashino, · 2017
Later among the works it cites.
“SEGAN: Speech enhancement generative adversarial network,”
S. Pascual, A. Bonafonte, and J. Serrá, · 2017
Later among the works it cites.
“Generative adversarial network-based postfilter for STFT spectrograms,”
T. Kaneko, S. Takaki, H. Kameoka, and J. Yamagishi, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Nakashika, T. Takiguchi, and Y. Ariki, · 2015
Cited alongside, same era.
“Autoencoding beyond pixels using a learned similarity metric,”
A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther, · 2015
Cited alongside, same era.
“Modeling and transforming speech using variational autoencoders,”
M. Blaauw and J. Bonada, · 2016
Cited alongside, same era.
“Voice conversion from non-parallel corpora using variational auto-encoder,”
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, · 2016
Cited alongside, same era.
“A KL divergence and DNN-based approach to voice conversion without parallel training sentences,”
F.-L. Xie, F. K. Soong, and H. Li, · 2016
Cited alongside, same era.
“WaveNet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Cited alongside, same era.
“WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,”
M. Morise, F. Yokomori, and K. Ozawa, · 2016
Cited alongside, same era.
“Unpaired image-to-image translation using cycle-consistent adversarial networks,”
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, · 2017
Later among the works it cites.
“Learning to discover cross-domain relations with generative adversarial networks,”
T. Kim, M. Cha, H. Kim, J. K. Lee, and J. Kim, · 2017
Later among the works it cites.
“DualGAN: Unsupervised dual learning for image-to-image translation,”
Z. Yi, H. Zhang, P. Tan, and M. Gong, · 2017
Later among the works it cites.
“StarGAN: Unified generative adversarial networks for multi-domain image-to-image translation,”
Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo, · 2017
Later among the works it cites.
“Parallel WaveNet: Fast high-fidelity speech synthesis,”
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. C. Cobo, F. Stimberg, N. Casagrande, D. G., S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, · 2017
Later among the works it cites.
“Language modeling with gated convolutional networks,”
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, · 2017
Later among the works it cites.
“Image-to-image translation with conditional adversarial networks,”
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, · 2017
Later among the works it cites.
“Non-parallel voice conversion using variational autoencoders conditioned by phonetic posteriorgrams and d-vectors,”
Y. Saito, Y. Ijima, K. Nishida, and S. Takamichi, · 2018
Closest in time.
“Statistical parametric speech synthesis incorporating generative adversarial networks,”
Y. Saito, S. Takamichi, and H. Saruwatari, · 2018
Closest in time.
K. Oyamada, H. Kameoka, T. Kaneko, K. Tanaka, N. Hojo, and H. Ando, · 2018
Closest in time.
“Deep clustering with gated convolutional networks,”
L. Li and H. Kameoka, · 2018
Closest in time.
“The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,”
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, · 2018
Closest in time.