Fetching the paper…
Reading the bibliography…
Although voice conversion (VC) algorithms have achieved remarkable success along with the development of machine learning, superior performance is still difficult to achieve when using nonparallel data.
“Voice conversion through vector quantization,”
M. Abe, S. Nakamura, K. Shikano, and H. Kuwabara, · 1988
Earlier work this paper cites.
“Voice conversion,”
D.G. Childers, K. Wu, D.M. Hicks, and B. Yegnanarayana, · 1989
Earlier work this paper cites.
“Voice conversion through vector quantization,”
M. Abe, S. Nakamura, K. Shikano, and H. Kuwabara, · 1990
Earlier work this paper cites.
“Transformation of formants for voice conversion using artificial neural networks,”
M. Narendranath, H. Murthy, S. Rajendran, and B. Yegnanarayana, · 1995
Earlier work this paper cites.
“Continuous probabilistic transform for voice conversion,”
Y. Stylianou, O. Cappé, and E. Moulines, · 1998
Earlier work this paper cites.
“Speech parameter generation algorithms for HMM-based speech synthesis,”
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, · 2000
Earlier work this paper cites.
“Subband based voice conversion,”
O. Turk and L. Arslan, · 2002
Earlier work this paper cites.
“Voice conversion for unknown speakers,”
H. Ye and S. Young, · 2004
Earlier work this paper cites.
“Framewise phoneme classification with bidirectional lstm and other neural network architectures,”
A. Graves and J. Schmidhuber, · 2005
Earlier work this paper cites.
“Incorporating a mixed excitation model and postfilter into HMM-based text-to-speech synthesis,”
T. Yoshimura, K. Tokuda, T. Masuko, T. Kobayashi, and T. Kitamura, · 2005
Earlier work this paper cites.
“Maximum likelihood voice conversion based on GMM with straight mixed excitation,”
Y. Ohtani, T. Toda, H. Saruwatari, and K. Shikano, · 2006
Earlier work this paper cites.
“Improving the intelligibility of dysarthric speech,”
A. Kain, J. Hosom, X. Niu, J. Santen, M. Fried-Oken, and J. Staehely, · 2007
Cited alongside, same era.
“Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
T. Toda, A. Black, and K. Tokuda, · 2007
Cited alongside, same era.
“Speech signal processing toolkit (SPTK),”
SPTK Working Group et al., · 2009
Cited alongside, same era.
“Spectral mapping using artificial neural networks for voice conversion,”
S. Desai, A. Black, B. Yegnanarayana, and K. Prahallad, · 2010
Cited alongside, same era.
“INCA algorithm for training voice conversion systems from nonparallel corpora,”
D. Erro, A. Moreno, and A. Bonafonte, · 2010
Cited alongside, same era.
“Feedback utterances for computer-adied language learning using accent reduction and voice conversion method,”
Software available from tensorflow.org
“TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, · 2015
Later among the works it cites.
“Non-parallel training in voice conversion using an adaptive restricted Boltzmann machine,”
T. Nakashika, T. Takiguchi, and Y. Minami, · 2016
Later among the works it cites.
“Merlin: An open source neural network speech synthesis system,”
Z. Wu, O. Watts, and S. King, · 2016
Later among the works it cites.
“World: A vocoder-based high-quality speech synthesis system for real-time applications,”
M. Morise, F. Yokomori, and K. Ozawa, · 2016
Later among the works it cites.
“An overview of voice conversion systems,”
S. Mohammadi and A. Kain, · 2017
Later among the works it cites.
“Unpaired image-to-image translation using cycle-consistent adversarial networks,”
J. Zhu, T. Park, P. Isola, and A. Efros, · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Zhao, S. Koh, S. Yann, and K. Luke, · 2013
Cited alongside, same era.
“Non-parallel training for voice conversion based on adaptation method,”
P. Song, W. Zheng, and L. Zhao, · 2013
Cited alongside, same era.
“Voice timbre control based on perceived age in singing voice conversion,”
K. Kobayashi, T. Toda, H. Doi, T. Nakano, M. Goto, G. Neubig, S. Sakti, and S. Nakamura, · 2014
Cited alongside, same era.
“Generative adversarial nets,”
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, · 2014
Cited alongside, same era.
“Non-parallel voice conversion using joint optimization of alignment by temporal context and spectral distortion,”
H. Benisty, D. Malah, and K. Crammer, · 2014
Cited alongside, same era.
“Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,”
L. Sun, S. Kang, K. Li, and H. Meng, · 2015
Cited alongside, same era.
“ALAGIN Japanese Speech Database,” http://shachi.org/resources/4255?ln=eng
Cited in the paper.
Later among the works it cites.
“Sequence-to-sequence voice conversion with similarity metric learned using generative adversarial networks,”
T. Kaneko, H. Kameoka, K. Hiramatsu, and K. Kashino, · 2017
Later among the works it cites.
“Voice conversion from unaligned corpora using variational autoencoding Wasserstein generative adversarial networks,”
C. Hsu, H. Hwang, Y. Wu, Y. Tsao, and H. Wang, · 2017
Later among the works it cites.
“Wasserstein generative adversarial networks,”
M. Arjovsky, S. Chintala, and L. Bottou, · 2017
Later among the works it cites.
“Least squares generative adversarial networks,”
X. Mao, Q. Li, H. Xie, R. YK. Lau, Z. Wang, and S. P. Smolley, · 2017
Later among the works it cites.