Fetching the paper…
Reading the bibliography…
So far, many of the deep learning approaches for voice conversion produce good quality speech by using a large amount of training data.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” Communications, Computers and Signal Processing
1993
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE transactions on neural networks
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation
1997
Earlier work this paper cites.
A. Kain and M. W. Macon, “Spectral voice conversion for text-to-speech synthesis,” in Acoustics, Speech and Signal Processing, 1998. Proceedings of the 1998 IEEE International Conference on
1998
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. de Cheveign´e, “Restructuring speech representations using a pitch-adaptive time–frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds,” Speech communication
1999
Earlier work this paper cites.
J. Kominek and A. W. Black, “The cmu arctic speech databases,” in Fifth ISCA Workshop on Speech Synthesis
2004
Earlier work this paper cites.
A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional lstm and other neural network architectures,” Neural Networks
2005
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,” IEEE Transactions on Audio, Speech and Language Processing
2007
Earlier work this paper cites.
M. Wöllmer, A. Metallinou, F. Eyben, B. Schuller, and S. Narayanan, “Context-sensitive multimodal emotion recognition from speech and facial expression using bidirectional lstm modeling,” in Proc. INTERSPEECH 2010, Makuhari, Japan
2010
Earlier work this paper cites.
D. Povey, A. Ghoshal, N. Goel, M. Hannemann, Y. Qian, P. Schwarz, J. Silovsk, and P. Motl, “The Kaldi Speech Recognition Toolkit,” In IEEE ASRU
2011
Earlier work this paper cites.
T. Toda, M. Nakagiri, and K. Shikano, “Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing
2012
Earlier work this paper cites.
K. Nakamura, T. Toda, H. Saruwatari, and K. Shikano, “Speaking-aid systems using gmm-based voice conversion for electrolaryngeal speech,” Speech Communication
2012
Earlier work this paper cites.
E. Helander, H. Silen, T. Virtanen, and M. Gabbouj, “Voice conversion using dynamic kernel partial least squares regression,” IEEE Transactions on Audio, Speech, and Language Processing
2012
Earlier work this paper cites.
R. Takashima, T. Takiguchi, and Y. Ariki, “Exemplar-based voice conversion using sparse representation in noisy environments,” IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences
2013
Cited alongside, same era.
T. Nakashika, R. Takashima, T. Takiguchi, and Y. Ariki, “Voice conversion in high-order eigen space using deep belief nets,” In INTERSPEECH
2013
Cited alongside, same era.
A. Graves, N. Jaitly, and A.-r. Mohamed, “Hybrid speech recognition with deep bidirectional lstm,” in Automatic Speech Recognition and Understanding (ASRU), 2013 IEEE Workshop on
2013
Cited alongside, same era.
M. Wöllmer, Z. Zhang, F. Weninger, B. Schuller, and G. Rigoll, “Feature enhancement by bidirectional lstm networks for conversational speech recognition in highly non-stationary noise,” in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on
2013
Cited alongside, same era.
L. Sun, S. Kang, K. Li, and H. Meng, “Voice conversion using deep bidirectional Long Short-Term Memory based Recurrent Neural Networks,” In ICASSP
2015
Later among the works it cites.
L. He, D. Jiang, L. Yang, E. Pei, P. Wu, and H. Sahli, “Multimodal affective dimension prediction using deep bidirectional long short-term memory recurrent neural networks,” in Proceedings of the 5th International Workshop on Audio/Visual Emotion Challenge
2015
Later among the works it cites.
F. Weninger, J. Bergmann, and B. Schuller, “Introducing currennt: The munich open-source cuda recurrent neural network toolkit,” The Journal of Machine Learning Research
2015
Later among the works it cites.
X. Tian, S. W. Lee, Z. Wu, E. S. Chng, S. Member, and H. Li, “An Exemplar-based Approach to Frequency Warping for Voice Conversion,” pp. 1–10, 2016
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Wu, T. Virtanen, E. S. Chng, and H. Li, “Exemplar-based sparse representation with residual compensation for voice conversion,” IEEE/ACM Transactions on Audio, Speech and Language Processing
2014
Cited alongside, same era.
X. Tian, Z. Wu, S. W. Lee, and E. S. Chng, “Correlation-based frequency warping for voice conversion,” in Chinese Spoken Language Processing (ISCSLP), 2014 9th International Symposium on
2014
Cited alongside, same era.
L.-h. Chen, Z.-h. Ling, L.-j. Liu, and L.-r. Dai, “Voice Conversion Using Deep Neural Networks With Layer-Wise Generative Training,” IEEE Transactions on Audio, Speech and Language Processing
2014
Cited alongside, same era.
S. H. Mohammadi and A. Kain, “Voice conversion using deep neural networks with speaker-independent pre-training,” in Spoken Language Technology Workshop (SLT), 2014 IEEE
2014
Cited alongside, same era.
T. Nakashika, T. Takiguchi, and Y. Ariki, “High-order sequence modeling using speaker-dependent recurrent temporal restricted Boltzmann machines for voice conversion,” In Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
2014
Cited alongside, same era.
Y. Fan, Y. Qian, F.-L. Xie, and F. K. Soong, “Tts synthesis with bidirectional lstm based recurrent neural networks,” in Fifteenth Annual Conference of the International Speech Communication Association
2014
Cited alongside, same era.
S. Takamichi, T. Toda, A. W. Black, and S. Nakamura, “Modulation spectrum-constrained trajectory training algorithm for gmm-based voice conversion,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on
2015
Cited alongside, same era.
X. Tian, Z. Wu, S. W. Lee, N. Q. Hy, M. Dong, and E. S. Chng, “System fusion for high-performance voice conversion,” In Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
2015
Cited alongside, same era.
S. Takamichi, T. Toda, A. W. Black, G. Neubig, S. Sakti, and S. Nakamura, “Postfilters to modify the modulation spectrum for statistical parametric speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing
2016
Later among the works it cites.
N. Xu, X. Yao, A. Jiang, X. Liu, and J. Bao, “High quality voice conversion by post-filtering the outputs of gaussian processes,” in 2016 24th European Signal Processing Conference (EUSIPCO)
2016
Later among the works it cites.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from non-parallel corpora using variational auto-encoder,” in Signal and Information Processing Association Annual Summit and Conference (APSIPA), 2016 Asia-Pacific
2016
Later among the works it cites.
K. Tanaka, S. Hara, M. Abe, M. Sato, and S. Minagi, “Speaker dependent approach for enhancing a glossectomy patient’s speech via gmm-based voice conversion,” Proc. Interspeech 2017
2017
Later among the works it cites.
B. Çişman, H. Li, and K. C. Tan, “Sparse representation of phonetic features for voice conversion with and without parallel data,” in Automatic Speech Recognition and Understanding Workshop (ASRU), 2017 IEEE
2017
Later among the works it cites.
B. Sisman, H. Li, and K. C. Tan, “Transformation of prosody in voice conversion,” in 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Zeyer, P. Doetsch, P. Voigtlaender, R. Schlüter, and H. Ney, “A comprehensive study of deep bidirectional lstm rnns for acoustic modeling in speech recognition,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on
2017
Later among the works it cites.
J. Wu, D. Huang, L. Xie, and H. Li, “Denoising recurrent neural network for deep bidirectional lstm based voice conversion,” Proc. Interspeech 2017
2017
Later among the works it cites.