Fetching the paper…
Reading the bibliography…
We present the Voice Conversion Challenge 2018, designed as a follow up to the 2016 edition with the aim of providing a common framework for evaluating and comparing different state-of-the-art voice conversion (VC) systems.
D. W. Griffin and J. S. Lim, “Signal estimation from modified short-time Fourier transform,”
1984
Earlier work this paper cites.
M. Abe, S. Nakamura, and K. Shikano, “Voice conversion through vector quantization,”
1990
Earlier work this paper cites.
Y. Stylianou, O. Cappé, and E. Moulines, “Continuous probabilistic transform for voice conversion,”
1998
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. de Cheveigné, “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based
1999
Earlier work this paper cites.
Y. Stylianou, “Applying the harmonic plus noise model in concatenative speech synthesis,”
2001
Earlier work this paper cites.
A. Kain and M. Macon, “Design and evaluation of a voice conversion algorithm based on spectral envelope mapping and residual prediction,” in
2001
Earlier work this paper cites.
D. Sündermann, H. Höge, A. Bonafonte, H. Ney, A. Black, and S. Narayanan, “Text-independent voice conversion based on unit selection,” in
2006
Earlier work this paper cites.
A. Mouchtaris, J. Van der Spiegel, and P. Mueller, “Nonparallel training for voice conversion based on a parameter adaptation approach,”
2006
Earlier work this paper cites.
H. Ye and S. Young, “Quality-enhanced voice morphing using maximum likelihood transformations,”
2006
Earlier work this paper cites.
A. Kain, J. Hosom, X. Niu, J. van Santen, M. Fried-Oken, and J. Staehely, “Improving the intelligibility of dysarthric speech,”
2007
Earlier work this paper cites.
T. Toda, A. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
2007
Earlier work this paper cites.
D. Felps, H. Bortfeld, and R. Gutierrez-Osuna, “Foreign accent conversion in computer assisted pronunciation training,”
2009
Earlier work this paper cites.
O. Türk and M. Schröder, “Evaluation of expressive speech synthesis with voice conversion and copy resynthesis techniques,”
2010
Earlier work this paper cites.
F. Villavicencio and J. Bonada, “Applying voice conversion to concatenative singing-voice synthesis,” in
2010
Earlier work this paper cites.
N. Pilkington, H. Zen, and M. Gales, “Gaussian process experts for voice conversion,” in
2011
Cited alongside, same era.
T. Toda, M. Nakagiri, and K. Shikano, “Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,”
2012
Cited alongside, same era.
E. Helander, H. Silé, T. Virtanen, and M. M. Gabbouj, “Voice conversion using dynamic kernel partial least squares regression,”
2012
Cited alongside, same era.
R. Takashima, T. Takiguchi, and Y. Ariki, “Exemplar-based voice conversion using sparse representation in noisy environments,”
2013
Cited alongside, same era.
T. Nakashika, T. Takiguchi, and Y. Ariki, “Voice conversion based on speaker-dependent restricted boltzmann machines,”
2014
Cited alongside, same era.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in
2016
Later among the works it cites.
K. Kobayashi, T. Toda, and S. Nakamura, “
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
Z. Wu, O. Watts, and S. King, “Merlin: An open source neural network speech synthesis system,” in
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L.-H. Chen, Z.-H. Ling, L.-J. Liu, and L.-R. Dai, “Voice conversion using deep neural networks with layer-wise generative training,”
2014
Cited alongside, same era.
Z. Wu, T. Virtanen, E. Chng, and H. Li, “Exemplar-based sparse representation with residual compensation for voice conversion,”
2014
Cited alongside, same era.
N. Xu, Y. Tang, J. Bao, A. Jiang, X. Liu, and Z. Yang, “Voice conversion based on gaussian processes by coherent and asymmetric training with limited training data,”
2014
Cited alongside, same era.
L. Sun, S. Kang, K. Li, and H. Meng, “Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,” in
2015
Cited alongside, same era.
G. J. Mysore, “Can we automatically transform speech recorded on common consumer devices in real-world environments into professional production quality speech? – a dataset, insights, and challenges,”
2015
Cited alongside, same era.
Z. Wu, P. L. D. Leon, C. Demiroglu, A. Khodabakhsh, S. King, Z. H. Ling, D. Saito, B. Stewart, T. Toda, M. Wester, and J. Yamagishi, “Anti-spoofing for text-independent speaker verification: An initial database, comparison of countermeasures, and human performance,”
2016
Cited alongside, same era.
T. Toda, L.-H. Chen, D. Saito, F. Villavicencio, M. Wester, Z. Wu, and J. Yamagishi, “The voice conversion challenge 2016,” in
2016
Cited alongside, same era.
Z. Wu, J. Yamagishi, T. Kinnunen, C. Hanilçi, M. Sahidullah, A. Sizov, N. Evans, M. Todisco, and H. Delgado, “Asvspoof: The automatic speaker verification spoofing and countermeasures challenge,”
2017
Later among the works it cites.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from unaligned corpora using variational autoencoding Wasserstein generative adversarial networks,” in
2017
Later among the works it cites.
Y. Saito, S. Takamichi, and H. Saruwatari, “Voice conversion using input-to-output highway networks,”
2017
Later among the works it cites.
S. Mohammadi and A. Kain, “An overview of voice conversion systems,”
2017
Later among the works it cites.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder,”
2017
Later among the works it cites.
Y.-J. Hu, C. Ding, L.-J. Liu, Z.-H. Ling, and L.-R. Dai, “The USTC system for Blizzard Challenge 2017.” in
2017
Later among the works it cites.
T. Kinnunen, J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, and Z. Ling, “A spoofing benchmark for the 2018 voice conversion challenge: Leveraging from spoofing countermeasures for speech quality assessment,” in
2018
Closest in time.
K. Kobayashi and T. Toda, “sprocket: Open-source voice conversion software,” in
2018
Closest in time.