Fetching the paper…
Reading the bibliography…
This paper proposes a voice conversion (VC) method using sequence-to-sequence (seq2seq or S2S) learning, which flexibly converts not only the voice characteristics but also the pitch contour and duration of input speech.
T. Fukada, K. Tokuda, T. Kobayashi, and S. Imai, “An adaptive algorithm for mel-cepstral analysis of speech,” in Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 1992, pp. 137–140
1992
Earlier work this paper cites.
A. Kain and M. W. Macon, “Spectral voice conversion for text-to-speech synthesis,” in Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 1998, pp. 285–288
1998
Earlier work this paper cites.
Y. Stylianou, O. Cappé, and E. Moulines, “Continuous probabilistic transform for voice conversion,” IEEE Trans. SAP , vol. 6, no. 2, pp. 131–142, 1998
1998
Earlier work this paper cites.
D. J. Hermes, “Measuring the perceptual similarity of pitch contours,” J. Speech Lang. Hear. Res. , vol. 41, no. 1, pp. 73–82, 1998
1998
Earlier work this paper cites.
J. Kominek and A. W. Black, “The CMU Arctic speech databases,” in Proc. ISCA Speech Synthesis Workshop (SSW) , 2004, pp. 223–224
2004
Earlier work this paper cites.
A. B. Kain, J.-P. Hosom, X. Niu, J. P. van Santen, M. Fried-Oken, and J. Staehely, “Improving the intelligibility of dysarthric speech,” Speech Communication , vol. 49, no. 9, pp. 743–759, 2007
2007
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 8, pp. 2222–2235, 2007
2007
Earlier work this paper cites.
Z. Inanoglu and S. Young, “Data-driven emotion conversion in spoken English,” Speech Communication , vol. 51, no. 3, pp. 268–283, 2009
2009
Earlier work this paper cites.
D. Felps, H. Bortfeld, and R. Gutierrez-Osuna, “Foreign accent conversion in computer assisted pronunciation training,” Speech Communication , vol. 51, no. 10, pp. 920–932, 2009
2009
Earlier work this paper cites.
O. Türk and M. Schröder, “Evaluation of expressive speech synthesis with voice conversion and copy resynthesis techniques,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 5, pp. 965–973, 2010
2010
Earlier work this paper cites.
E. Helander, T. Virtanen, J. Nurminen, and M. Gabbouj, “Voice conversion using partial least squares regression,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 5, pp. 912–921, 2010
2010
Earlier work this paper cites.
S. Desai, A. W. Black, B. Yegnanarayana, and K. Prahallad, “Spectral mapping using artificial neural networks for voice conversion,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 5, pp. 954–964, 2010
2010
Earlier work this paper cites.
K. Nakamura, T. Toda, H. Saruwatari, and K. Shikano, “Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speech,” Speech Communication , vol. 54, no. 1, pp. 134–146, 2012
2012
Earlier work this paper cites.
T. Toda, M. Nakagiri, and K. Shikano, “Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 9, pp. 2505–2517, 2012
2012
Earlier work this paper cites.
S. H. Mohammadi and A. Kain, “Voice conversion using deep neural networks with speaker-independent pre-training,” in Proc. IEEE Spoken Language Technology Workshop (SLT) , 2014, pp. 19–23
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Adv. Neural Information Processing Systems (NIPS) , 2014, pp. 3104–3112
2014
Earlier work this paper cites.
L. Sun, S. Kang, K. Li, and H. Meng, “Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,” in Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2015, pp. 4869–4873
2015
Earlier work this paper cites.
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Adv. Neural Information Processing Systems (NIPS) , 2015, pp. 577–585
2015
Earlier work this paper cites.
M.-T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” in Proc. EMNLP , 2015
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: accelerating deep network training by reducing internal covariate shift,” in Proc. International Conference on Machine Learning (ICML) , 2015, pp. 448–456
2015
Earlier work this paper cites.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from non-parallel corpora using variational auto-encoder,” in Proc. Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , 2016, pp. 1–6
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,” IEICE Transactions on Information and Systems , vol. E99-D, no. 7, pp. 1877–1884, 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
H. Tachibana, K. Uenoyama, and S. Aihara, “Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention,” in Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2018, pp. 4784–4788
2018
Closest in time.
W. Ping, K. Peng, A. Gibiansky, S. O. Arık, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep Voice 3: Scaling text-to-speech with convolutional sequence learning,” in Proc. International Conference on Learning Representations (ICLR) , 2018
2018
Closest in time.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural tts synthesis by conditioning WaveNet on mel spectrogram predictions,” in Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2018, pp. 4779–4783
2018
Closest in time.
H. Kameoka, K. Tanaka, T. Kaneko, and N. Hojo, “ConvS2S-VC: Fully convolutional sequence-to-sequence voice conversion,” arXiv , Nov. 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in Adv. Neural Information Processing Systems (NIPS) , 2016, pp. 901–909
2016
Cited alongside, same era.
Y. Saito, S. Takamichi, and H. Saruwatari, “Voice conversion using input-to-output highway networks,” IEICE Trans Inf. Syst. , vol. E100-D, no. 8, pp. 1925–1928, 2017
2017
Cited alongside, same era.
T. Kaneko, H. Kameoka, K. Hiramatsu, and K. Kashino, “Sequence-to-sequence voice conversion with similarity metric learned using generative adversarial networks,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2017, pp. 1283–1287
2017
Cited alongside, same era.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from unaligned corpora using variational autoencoding Wasserstein generative adversarial networks,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2017, pp. 3364–3368
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2017, pp. 4006–4010
2017
Cited alongside, same era.
S. O. Arık, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, S. Sengupta, and M. Shoeybi, “Deep voice: Real-time neural text-to-speech,” in Proc. International Conference on Machine Learning (ICML) , 2017
2017
Cited alongside, same era.
S. O. Arık, G. Diamos, A. Gibiansky, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, “Deep voice 2: Multi-speaker neural text-to-speech,” in Proc. Neural Information Processing Systems (NIPS) , 2017
2017
Cited alongside, same era.
2018
Closest in time.
2018
Closest in time.
A. Haque, M. Guo, and P. Verma, “Conditional end-to-end audio transforms,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2018, pp. 2295–2299
2018
Closest in time.
2018
Closest in time.
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “FFTNet: A real-time speaker-dependent neural vocoder,” in Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2018, pp. 2251–2255
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
K. Tanaka, T. Kaneko, N. Hojo, and H. Kameoka, “Synthetic-to-natural speech waveform conversion using cycle-consistent adversarial networks,” in Proc. IEEE Spoken Language Technology Workshop (SLT) , 2018, pp. 632–639
2018
Closest in time.
K. Kobayashi and T. Toda, “sprocket: Open-source voice conversion software,” in Proc. Odyssey , 2018, pp. 203–210
2018
Closest in time.
2018
Closest in time.
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “ACVAE-VC: Non-parallel voice conversion with auxiliary classifier variational autoencoder,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 9, pp. 1432–1443, 2019
2019
Closest in time.
K. Tanaka, H. Kameoka, T. Kaneko, and N. Hojo, “AttS2S-VC: Sequence-to-sequence voice conversion with attention and context preservation mechanisms,” in Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2019, pp. 6805–6809
2019
Closest in time.
M. Zhang, X. Wang, F. Fang, H. Li, and J. Yamagishi, “Joint training framework for text-to-speech and voice conversion using multi-source Tacotron and WaveNet,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2019, pp. 1298–1302
2019
Closest in time.
F. Biadsy, R. J. Weiss, P. J. Moreno, D. Kanevsky, and Y. Jia, “Parrotron: An end-to-end speech-to-speech conversion model and its applications to hearing-impaired speech and speech separation,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2019, pp. 4115–4119
2019
Closest in time.
2019
Closest in time.
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “StarGAN-VC2: Rethinking conditional methods for stargan-based voice conversion,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2019, pp. 679–683
2019
Closest in time.
https://github.com/k2kobayashi/sprocket, (Accessed on 01/28/2019)
2019
Closest in time.