Fetching the paper…
Reading the bibliography…
We propose a parallel-data-free voice-conversion (VC) method that can learn a mapping from source to target speech without relying on parallel data.
“Spectral voice conversion for text-to-speech synthesis,”
A. Kain and M. W. Macon, · 1998
Earlier work this paper cites.
“Continuous probabilistic transform for voice conversion,”
Y. Stylianou, O. Cappé, and E. Moulines, · 1998
Earlier work this paper cites.
“Nonparallel training for voice conversion based on a parameter adaptation approach,”
A. Mouchtaris, J. Van der Spiegel, and P. Mueller, · 2006
Earlier work this paper cites.
“MAP-based adaptation for speech conversion using adaptation data selection and non-parallel training,”
C.-H. Lee and C.-H. Wu, · 2006
Earlier work this paper cites.
“Eigenvoice conversion based on Gaussian mixture model,”
T. Toda, Y. Ohtani, and K. Shikano, · 2006
Earlier work this paper cites.
“Maximum likelihood voice conversion based on GMM with STRAIGHT mixed excitation,”
Y. Ohtani, T. Toda, H. Saruwatari, and K. Shikano, · 2006
Earlier work this paper cites.
“Improving the intelligibility of dysarthric speech,”
A. B. Kain, J.-P. Hosom, X. Niu, J. P. van Santen, M. Fried-Oken, and J. Staehely, · 2007
Earlier work this paper cites.
“Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
T. Toda, A. W. Black, and K. Tokuda, · 2007
Earlier work this paper cites.
“High quality voice conversion through phoneme-based linear mapping functions with STRAIGHT for Mandarin,”
K. Liu, J. Zhang, and Y. Yan, · 2007
Earlier work this paper cites.
“On the impact of alignment on voice conversion performance,”
E. Helander, J. Schwarz, J. Nurminen, H. Silen, and M. Gabbouj, · 2008
Earlier work this paper cites.
“Text-independent voice conversion based on state mapped codebook,”
M. Zhang, J. Tao, J. Tian, and X. Wang, · 2008
Earlier work this paper cites.
“Data-driven emotion conversion in spoken English,”
Z. Inanoglu and S. Young, · 2009
Earlier work this paper cites.
“Evaluation of expressive speech synthesis with voice conversion and copy resynthesis techniques,”
O. Türk and M. Schröder, · 2010
Earlier work this paper cites.
“Voice conversion using partial least squares regression,”
E. Helander, T. Virtanen, J. Nurminen, and M. Gabbouj, · 2010
Earlier work this paper cites.
“Spectral mapping using artificial neural networks for voice conversion,”
S. Desai, A. W. Black, B. Yegnanarayana, and K. Prahallad, · 2010
Earlier work this paper cites.
“Rectified linear units improve restricted Boltzmann machines,”
V. Nair and G. E. Hinton, · 2010
Earlier work this paper cites.
“One-to-many voice conversion based on tensor representation of speaker space,”
D. Saito, K. Yamamoto, N. Minematsu, and K. Hirose, · 2011
Earlier work this paper cites.
“Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speech,”
K. Nakamura, T. Toda, H. Saruwatari, and K. Shikano, · 2012
Earlier work this paper cites.
“Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,”
T. Toda, M. Nakagiri, and K. Shikano, · 2012
Earlier work this paper cites.
“Exampler-based voice conversion using sparse representation in noisy environments,”
R. Takashima, T. Takiguchi, and Y. Ariki, · 2013
Cited alongside, same era.
“Rectifier nonlinearities improve neural network acoustic models,”
A. Maas, A. Y. Hannun, and A. Y. Ng, · 2013
Cited alongside, same era.
“Voice conversion using deep neural networks with layer-wise generative training,”
L.-H. Chen, Z.-H. Ling, L.-J. Liu, and L.-R. Dai, · 2014
Cited alongside, same era.
“Voice conversion based on speaker-dependent restricted Boltzmann machines,”
T. Nakashika, T. Takiguchi, and Y. Ariki, · 2014
Cited alongside, same era.
“Voice conversion using deep neural networks with speaker-independent pre-training,”
S. H. Mohammadi and A. Kain, · 2014
Cited alongside, same era.
“High-order sequence modeling using speaker-dependent recurrent temporal restricted Boltzmann machines for voice conversion,”
“Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network,”
W. Shi, J. Caballero, F. Huszár, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, and Z. Wang, · 2016
Later among the works it cites.
“Perceptual losses for real-time style transfer and super-resolution,”
J. Johnson, A. Alahi, and L. Fei-Fei, · 2016
Later among the works it cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Later among the works it cites.
“Instance normalization: The missing ingredient for fast stylization,”
D. Ulyanov, A. Vedaldi, and V. Lempitsky, · 2016
Later among the works it cites.
“Least squares generative adversarial networks,”
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. P. Smolley, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Nakashika, T. Takiguchi, and Y. Ariki, · 2014
Cited alongside, same era.
“Exemplar-based sparse representation with residual compensation for voice conversion,”
Z. Wu, T. Virtanen, E. S. Chng, and H. Li, · 2014
Cited alongside, same era.
“Generative adversarial nets,”
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, · 2014
Cited alongside, same era.
“A postfilter to modify the modulation spectrum in HMM-based speech synthesis,”
S. Takamichi, T. Toda, G. Neubig, S. Sakti, and S. Nakamura, · 2014
Cited alongside, same era.
“Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,”
L. Sun, S. Kang, K. Li, and H. Meng, · 2015
Cited alongside, same era.
“Fully convolutional networks for semantic segmentation,”
J. Long, E. Shelhamer, and T. Darrell, · 2015
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
D. Kingma and J. Ba, · 2015
Cited alongside, same era.
“Multidimensional scaling of systems in the Voice Conversion Challenge 2016,”
M. Wester, Z. Wu, and J. Yamagishi, · 2016
Later among the works it cites.
“Analysis of the Voice Conversion Challenge 2016 evaluation results,”
M. Wester, Z. Wu, and J. Yamagishi, · 2016
Later among the works it cites.
“F0 transformation techniques for statistical voice conversion with direct waveform modification with spectral differential,”
K. Kobayashi, T. Toda, and S. Nakamura, · 2016
Later among the works it cites.
“Sequence-to-sequence voice conversion with similarity metric learned using generative adversarial networks,”
T. Kaneko, H. Kameoka, K. Hiramatsu, and K. Kashino, · 2017
Closest in time.
“Unpaired image-to-image translation using cycle-consistent adversarial networks,”
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, · 2017
Closest in time.
“Learning to discover cross-domain relations with generative adversarial networks,”
T. Kim, M. Cha, H. Kim, J. K. Lee, and J. Kim, · 2017
Closest in time.
“DualGAN: Unsupervised dual learning for image-to-image translation,”
Z. Yi, H. Zhang, P. Tan, and M. Gong, · 2017
Closest in time.
“Language modeling with gated convolutional networks,”
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, · 2017
Closest in time.
“Unsupervised cross-domain image generation,”
Y. Taigman, A. Polyak, and L. Wolf, · 2017
Closest in time.
“Generative adversarial network-based postfilter for statistical parametric speech synthesis,”
T. Kaneko, H. Kameoka, N. Hojo, Y. Ijima, K. Hiramatsu, and K. Kashino, · 2017
Closest in time.
“Generative adversarial network-based postfilter for STFT spectrograms,”
T. Kaneko, S. Takaki, H. Kameoka, and J. Yamagishi, · 2017
Closest in time.
“Training algorithm to deceive anti-spoofing verification for DNN-based speech synthesis,”
Y. Saito, S. Takamichi, and H. Saruwatari, · 2017
Closest in time.
“Non-parallel voice conversion using i-vector PLDA: Towards unifying speaker verification and transformation,”
T. Kinnunen, L. Juvela, P. Alku, and J. Yamagishi, · 2017
Closest in time.
“Voice conversion from unaligned corpora using variational autoencoding Wasserstein generative adversarial networks,”
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, · 2017
Closest in time.