Fetching the paper…
Reading the bibliography…
This paper proposes a voice conversion (VC) method based on a sequence-to-sequence (S2S) learning framework, which enables simultaneous conversion of the voice characteristics, pitch contour, and duration of input speech.
A. Kurematsu, K. Takeda, Y. Sagisaka, S. Katagiri, H. Kuwabara, and K. Shikano, “ATR Japanese speech database as a tool of speech recognition and synthesis,” Speech Communication , vol. 9, no. 4, pp. 357–363, Aug. 1990
1990
Earlier work this paper cites.
T. Fukada, K. Tokuda, T. Kobayashi, and S. Imai, “An adaptive algorithm for mel-cepstral analysis of speech,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 1992, pp. 137–140
1992
Earlier work this paper cites.
A. Kain and M. W. Macon, “Spectral voice conversion for text-to-speech synthesis,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 1998, pp. 285–288
1998
Earlier work this paper cites.
Y. Stylianou, O. Cappé, and E. Moulines, “Continuous probabilistic transform for voice conversion,” IEEE Trans. SAP , vol. 6, no. 2, pp. 131–142, 1998
1998
Earlier work this paper cites.
D. J. Hermes, “Measuring the perceptual similarity of pitch contours,” J. Speech Lang. Hear. Res. , vol. 41, no. 1, pp. 73–82, 1998
1998
Earlier work this paper cites.
J. Kominek and A. W. Black, “The CMU Arctic speech databases,” in Proc. ISCA Speech Synthesis Workshop (SSW) , 2004, pp. 223–224
2004
Earlier work this paper cites.
A. B. Kain, J.-P. Hosom, X. Niu, J. P. van Santen, M. Fried-Oken, and J. Staehely, “Improving the intelligibility of dysarthric speech,” Speech Communication , vol. 49, no. 9, pp. 743–759, 2007
2007
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 8, pp. 2222–2235, 2007
2007
Earlier work this paper cites.
Z. Inanoglu and S. Young, “Data-driven emotion conversion in spoken English,” Speech Communication , vol. 51, no. 3, pp. 268–283, 2009
2009
Earlier work this paper cites.
D. Felps, H. Bortfeld, and R. Gutierrez-Osuna, “Foreign accent conversion in computer assisted pronunciation training,” Speech Communication , vol. 51, no. 10, pp. 920–932, 2009
2009
Earlier work this paper cites.
E. Helander, T. Virtanen, J. Nurminen, and M. Gabbouj, “Voice conversion using partial least squares regression,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 5, pp. 912–921, 2010
2010
Earlier work this paper cites.
S. Desai, A. W. Black, B. Yegnanarayana, and K. Prahallad, “Spectral mapping using artificial neural networks for voice conversion,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 5, pp. 954–964, 2010
2010
Earlier work this paper cites.
E. Helander, H. Silen, T. Virtanen, and M. Gabbouj, “Voice conversion using dynamic kernel partial least squares regression,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 3, pp. 806–817, 2011
2011
Earlier work this paper cites.
K. Nakamura, T. Toda, H. Saruwatari, and K. Shikano, “Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speech,” Speech Communication , vol. 54, no. 1, pp. 134–146, 2012
2012
Earlier work this paper cites.
T. Toda, M. Nakagiri, and K. Shikano, “Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 9, pp. 2505–2517, 2012
2012
Earlier work this paper cites.
R. Takashima, T. Takiguchi, and Y. Ariki, “Exemplar-based voice conversion in noisy environment,” in Proc. IEEE Spoken Language Technology Workshop (SLT) , 2012, pp. 313–317
2012
Earlier work this paper cites.
S. H. Mohammadi and A. Kain, “Voice conversion using deep neural networks with speaker-independent pre-training,” in Proc. IEEE Spoken Language Technology Workshop (SLT) , 2014, pp. 19–23
2014
Earlier work this paper cites.
G. Sanchez, H. Silen, J. Nurminen, and M. Gabbouj, “Hierarchical modeling of F0 contours for voice conversion,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2014, pp. 2318–2321
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Adv. Neural Information Processing Systems (NIPS) , 2014, pp. 3104–3112
2014
Earlier work this paper cites.
L. Sun, S. Kang, K. Li, and H. Meng, “Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 4869–4873
2015
Earlier work this paper cites.
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Adv. Neural Information Processing Systems (NIPS) , 2015, pp. 577–585
2015
Earlier work this paper cites.
M.-T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” in Proc. EMNLP , 2015
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
H. Ming, D. Huang, L. Xie, J. Wu, M. Dong, and H. Li, “Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2016, pp. 2453–2457
2016
Earlier work this paper cites.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from non-parallel corpora using variational auto-encoder,” in Proc. Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , 2016, pp. 1–6
2016
Earlier work this paper cites.
H. Ming, D. Huang, L. Xie, J. Wu, M. Dong, and H. Li, “Exemplar-based sparse representation of timbre and prosody for voice conversion,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016, pp. 5175–5179
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,” IEICE Transactions on Information and Systems , vol. E99-D, no. 7, pp. 1877–1884, 2016
2016
Cited alongside, same era.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in Adv. Neural Information Processing Systems (NIPS) , 2016, pp. 901–909
2016
Cited alongside, same era.
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “FFTNet: A real-time speaker-dependent neural vocoder,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 2251–2255
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Tian, S. W. Lee, Z. Wu, E. S. Chng, and H. Li, “An exemplar-based approach to frequency warping for voice conversion,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 10, pp. 1863–1876, 2017
2017
Cited alongside, same era.
B. Sisman, H. Li, and K. C. Tan, “Sparse representation of phonetic features for voice conversion with and without parallel data,” in Proc. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2017, pp. 677–684
2017
Cited alongside, same era.
——, “Voice conversion from unaligned corpora using variational autoencoding Wasserstein generative adversarial networks,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2017, pp. 3364–3368
2017
Cited alongside, same era.
Z. Luo, J. Chen, T. Takiguchi, and Y. Ariki, “Emotional voice conversion with adaptive scales F0 based on wavelet transform using limited amount of emotional data,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2017, p. 3399.3403
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2017, pp. 4006–4010
2017
Cited alongside, same era.
S. O. Arık, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, S. Sengupta, and M. Shoeybi, “Deep Voice: Real-time neural text-to-speech,” in Proc. International Conference on Machine Learning (ICML) , 2017
2017
Cited alongside, same era.
S. O. Arık, G. Diamos, A. Gibiansky, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, “Deep Voice 2: Multi-speaker neural text-to-speech,” in Adv. Neural Information Processing Systems (NIPS) , 2017
2017
Cited alongside, same era.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, “Char2Wav: End-to-end speech synthesis,” in Proc. International Conference on Learning Representations (ICLR) , 2017
2017
Cited alongside, same era.
K. Tanaka, T. Kaneko, N. Hojo, and H. Kameoka, “Synthetic-to-natural speech waveform conversion using cycle-consistent adversarial networks,” in Proc. IEEE Spoken Language Technology Workshop (SLT) , 2018, pp. 632–639
2018
Later among the works it cites.
K. Kobayashi and T. Toda, “sprocket: Open-source voice conversion software,” in Proc. Odyssey , 2018, pp. 203–210
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “ACVAE-VC: Non-parallel voice conversion with auxiliary classifier variational autoencoder,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 9, pp. 1432–1443, 2019
2019
Later among the works it cites.
P. L. Tobing, Y.-C. Wu, T. Hayashi, K. Kobayashi, and T. Toda, “Non-parallel voice conversion with cyclic variational autoencoder,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2019, pp. 674–678
2019
Later among the works it cites.
2019
Later among the works it cites.
B. Sisman, M. Zhang, and H. Li, “Group sparse representation with WaveNet vocoder adaptation for spectrum and prosody conversion,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 6, pp. 1085–1097, 2019
2019
Later among the works it cites.
M. Zhang, X. Wang, F. Fang, H. Li, and J. Yamagishi, “Joint training framework for text-to-speech and voice conversion using multi-source Tacotron and WaveNet,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2019, pp. 1298–1302
2019
Later among the works it cites.
F. Biadsy, R. J. Weiss, P. J. Moreno, D. Kanevsky, and Y. Jia, “Parrotron: An end-to-end speech-to-speech conversion model and its applications to hearing-impaired speech and speech separation,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2019, pp. 4115–4119
2019
Later among the works it cites.
K. Tanaka, H. Kameoka, T. Kaneko, and N. Hojo, “AttS2S-VC: Sequence-to-sequence voice conversion with attention and context preservation mechanisms,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6805–6809
2019
Later among the works it cites.
Q. Wang, B. Li, T. Xiao, J. Zhu, C. Li, D. F. Wong, and L. S. Chao, “Learning deep transformer models for machine translation,” in Proc. Annual Meeting of the Association for Computational Linguistics (ACL) , 2019, pp. 1810–1822
2019
Later among the works it cites.
2019
Later among the works it cites.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brebisson, Y. Bengio, and A. Courville, “MelGAN: Generative adversarial networks for conditional waveform synthesis,” in Adv. Neural Information Processing Systems (NeurIPS) , 2019, pp. 14 910–14 921
2019
Later among the works it cites.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-shot voice style transfer with only autoencoder loss,” in Proc. International Conference on Machine Learning (ICML) , 2019, pp. 5210–5219
2019
Later among the works it cites.
2020
Closest in time.
H. Kameoka, K. Tanaka, D. Kwaśny, T. Kaneko, and N. Hojo, “ConvS2S-VC: Fully convolutional sequence-to-sequence voice conversion,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 1849–1863, 2020
2020
Closest in time.
W.-C. Huang, T. Hayashi, Y.-C. Wu, H. Kameoka, and T. Toda, “Voice transformer network: Sequence-to-sequence voice conversion using transformer with text-to-speech pretraining,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2020
2020
Closest in time.
2020
Closest in time.
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T.-Y. Liu, “On layer normalization in the transformer architecture,” in Proc. International Conference on Machine Learning (ICML) , 2020, pp. 503–512
2020
Closest in time.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6199–6203
2020
Closest in time.