Fetching the paper…
Reading the bibliography…
This paper proposes a non-parallel many-to-many voice conversion (VC) method using a variant of the conditional variational autoencoder (VAE) called an auxiliary classifier VAE (ACVAE).
“An adaptive algorithm for mel-cepstral analysis of speech,”
T. Fukada, K. Tokuda, T. Kobayashi, and S. Imai, · 1992
Earlier work this paper cites.
“Spectral voice conversion for text-to-speech synthesis,”
A. Kain and M. W. Macon, · 1998
Earlier work this paper cites.
“Continuous probabilistic transform for voice conversion,”
Y. Stylianou, O. Cappé, and E. Moulines, · 1998
Earlier work this paper cites.
“The IM algorithm: A variational approach to information maximization,”
D. Barber and F. V. Agakov, · 2003
Earlier work this paper cites.
“Improving the intelligibility of dysarthric speech,”
A. B. Kain, J.-P. Hosom, X. Niu, J. P. van Santen, M. Fried-Oken, and J. Staehely, · 2007
Earlier work this paper cites.
“Voice conversion based on maximumlikelihood estimation of spectral parameter trajectory,”
T. Toda, A. W. Black, and K. Tokuda, · 2007
Earlier work this paper cites.
K. Liu, J. Zhang, and Y. Yan, · 2007
Earlier work this paper cites.
“Data-driven emotion conversion in spoken English,”
Z. Inanoglu and S. Young, · 2009
Earlier work this paper cites.
“Evaluation of expressive speech synthesis with voice conversion and copy resynthesis techniques,”
O. Türk and M. Schröder, · 2010
Earlier work this paper cites.
“Voice conversion using partial least squares regression,”
E. Helander, T. Virtanen, J. Nurminen, and M. Gabbouj, · 2010
Earlier work this paper cites.
“Spectral mapping using artificial neural networks for voice conversion,”
S. Desai, A. W. Black, B. Yegnanarayana, and K. Prahallad, · 2010
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, · 2011
Earlier work this paper cites.
“Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speech,”
K. Nakamura, T. Toda, H. Saruwatari, and K. Shikano, · 2012
Earlier work this paper cites.
“Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,”
T. Toda, M. Nakagiri, and K. Shikano, · 2012
Earlier work this paper cites.
“Exampler-based voice conversion using sparse representation in noisy environments,”
R. Takashima, T. Takiguchi, and Y. Ariki, · 2013
Cited alongside, same era.
“Voice conversion using deep neural networks with layer-wise generative training,”
L.-H. Chen, Z.-H. Ling, L.-J. Liu, and L.-R. Dai, · 2014
Cited alongside, same era.
“Voice conversion based on speaker-dependent restricted Boltzmann machines,”
T. Nakashika, T. Takiguchi, and Y. Ariki, · 2014
Cited alongside, same era.
“Voice conversion using deep neural networks with speaker-independent pre-training,”
S. H. Mohammadi and A. Kain, · 2014
Cited alongside, same era.
“High-order sequence modeling using speaker-dependent recurrent temporal restricted boltzmann machines for voice conversion,”
T. Nakashika, T. Takiguchi, and Y. Ariki, · 2014
Cited alongside, same era.
“InfoGAN: Interpretable representation learning by information maximizing generative adversarial nets,”
X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel, · 2016
Later among the works it cites.
“WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,”
M. Morise, F. Yokomori, and K. Ozawa, · 2016
Later among the works it cites.
“Sequence-to-sequence voice conversion with similarity metric learned using generative adversarial networks,”
T. Kaneko, H. Kameoka, K. Hiramatsu, and K. Kashino, · 2017
Later among the works it cites.
“Voice conversion from unaligned corpora using variational autoencoding Wasserstein generative adversarial networks,”
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, · 2017
Later among the works it cites.
“Non-parallel voice conversion using i-vector PLDA: Towards unifying speaker verification and transformation,”
T. Kinnunen, L. Juvela, P. Alku, and J. Yamagishi, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Wu, T. Virtanen, E. S. Chng, and H. Li, · 2014
Cited alongside, same era.
“Auto-encoding variational Bayes,”
D. P. Kingma and M. Welling, · 2014
Cited alongside, same era.
“Semi-supervised learning with deep generative models,”
D. P. Kingma and D. J. Rezendey, S. Mohamedy, and M. Welling, · 2014
Cited alongside, same era.
“Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,”
L. Sun, S. Kang, K. Li, and H. Meng, · 2015
Cited alongside, same era.
“Autoencoding beyond pixels using a learned similarity metric,”
A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther, · 2015
Cited alongside, same era.
“Modeling and transforming speech using variational autoencoders,”
M. Blaauw and J. Bonada, · 2016
Cited alongside, same era.
“Voice conversion from non-parallel corpora using variational auto-encoder,”
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, · 2016
Cited alongside, same era.
“Conditional image synthesis with auxiliary classifier GANs,”
A. Odena, C. Olah, and J. Shlens, · 2017
Later among the works it cites.
“StarGAN: Unified generative adversarial networks for multi-domain image-to-image translation,”
Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo, · 2017
Later among the works it cites.
“Language modeling with gated convolutional networks,”
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, · 2017
Later among the works it cites.
“Parallel-data-free voice conversion using cycle-consistent adversarial networks,”
T. Kaneko and H. Kameoka, · 2017
Later among the works it cites.
“Non-parallel voice conversion using variational autoencoders conditioned by phonetic posteriorgrams and d-vectors,”
Y. Saito, Y. Ijima, K. Nishida, and S. Takamichi, · 2018
Closest in time.
“StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks,”
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, · 2018
Closest in time.
“Deep clustering with gated convolutional networks,”
L. Li and H. Kameoka, · 2018
Closest in time.
“The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,”
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, · 2018
Closest in time.