Fetching the paper…
Reading the bibliography…
In a typical voice conversion system, vocoder is commonly used for speech-to-features analysis and features-to-speech synthesis.
H. Sakoe and S. Chiba, “Dynamic programming algorithm optimization for spoken word recognition,”
1978
Earlier work this paper cites.
D. B. Paul and J. M. Baker, “The design for the wall street journal-based CSR corpus,” in
1992
Earlier work this paper cites.
Y. Stylianou, O. Cappé, and E. Moulines, “Continuous probabilistic transform for voice conversion,”
1998
Earlier work this paper cites.
A. Kain and M. W. Macon, “Spectral voice conversion for text-to-speech synthesis,” in
1998
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. de Cheveigné, “Restructuring speech representations using a pitch-adaptive time–frequency smoothing and an instantaneous-frequency-based F0 extraction: Possible role of a repetitive structure in sounds,”
1999
Earlier work this paper cites.
J. Kominek and A. W. Black, “The CMU arctic speech databases,” in
2004
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
2007
Earlier work this paper cites.
S. Desai, E. V. Raghavendra, B. Yegnanarayana, A. W. Black, and K. Prahallad, “Voice conversion using artificial neural networks,” in
2009
Earlier work this paper cites.
D. Erro, A. Moreno, and A. Bonafonte, “Voice conversion based on weighted frequency warping,”
2010
Earlier work this paper cites.
H. Benisty and D. Malah, “Voice conversion using GMM with enhanced global variance,” in
2011
Earlier work this paper cites.
E. Godoy, O. Rosec, and T. Chonavel, “Voice conversion using dynamic frequency warping with amplitude scaling, for parallel or nonparallel corpora,”
2012
Earlier work this paper cites.
R. Takashima, T. Takiguchi, and Y. Ariki, “Exemplar-based voice conversion in noisy environment,” in
2012
Cited alongside, same era.
T. Toda, T. Muramatsu, and H. Banno, “Implementation of computationally efficient real-time voice conversion.” in
2012
Cited alongside, same era.
X. Tian, Z. Wu, S. W. Lee, and E. S. Chng, “Correlation-based frequency warping for voice conversion,” in
2014
Cited alongside, same era.
Z. Wu, T. Virtanen, E. S. Chng, and H. Li, “Exemplar-based sparse representation with residual compensation for voice conversion,”
2014
Cited alongside, same era.
L.-H. Chen, Z.-H. Ling, L.-J. Liu, and L.-R. Dai, “Voice conversion using deep neural networks with layer-wise generative training,”
2014
Cited alongside, same era.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from unaligned corpora using variational autoencoding wasserstein generative adversarial networks,”
2017
Later among the works it cites.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder,” in
2017
Later among the works it cites.
T. Hayashi, A. Tamamori, K. Kobayashi, K. Takeda, and T. Toda, “An investigation of multi-speaker training for WaveNet vocoder,” in
2017
Later among the works it cites.
K. Kobayashi, T. Hayashi, A. Tamamori, and T. Toda, “Statistical voice conversion with WaveNet-based waveform generation,” in
2017
Later among the works it cites.
Y. Saito, S. Takamichi, and H. Saruwatari, “Statistical parametric speech synthesis incorporating generative adversarial networks,”
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F.-L. Xie, Y. Qian, Y. Fan, F. K. Soong, and H. Li, “Sequence error (se) minimization training of neural network for voice conversion,” in
2014
Cited alongside, same era.
X. Tian, Z. Wu, S. W. Lee, N. Q. Hy, E. S. Chng, and M. Dong, “Sparse representation for frequency warping based voice conversion,” in
2015
Cited alongside, same era.
L. Sun, S. Kang, K. Li, and H. Meng, “Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,” in
2015
Cited alongside, same era.
M. Morise, F. Yokomori, and K. Ozawa, “World: a vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Cited alongside, same era.
A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, “WaveNet: A generative model for raw audio.” in
2016
Cited alongside, same era.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in
2016
Cited alongside, same era.
N. Adiga, V. Tsiaras, and Y. Stylianou, “On the use of WaveNet as a statistical vocoder,” in
2018
Later among the works it cites.
Y.-C. Wu, P. L. Tobing, T. Hayashi, K. Kobayashi, and T. Toda, “The nu non-parallel voice conversion system for the voice conversion challenge 2018,”
2018
Later among the works it cites.
L.-J. Liu, Z.-H. Ling, Y. Jiang, M. Zhou, and L.-R. Dai, “Wavenet vocoder with limited training data for voice conversion,”
2018
Later among the works it cites.
B. Sisman, M. Zhang, and H. Li, “A voice conversion framework with tandem feature sparse representation and speaker-adapted wavenet vocoder,”
2018
Later among the works it cites.
Y.-C. Wu, K. Kobayashi, T. Hayashi, P. L. Tobing, and T. Toda, “Collapsed speech segment detection and suppression for WaveNet vocoder,”
2018
Later among the works it cites.
X. Tian, J. Wang, H. Xu, E.-S. Chng, and H. Li, “Average modeling approach to voice conversion with non-parallel data,” in
2018
Later among the works it cites.