Fetching the paper…
Reading the bibliography…
This paper presents a method of sequence-to-sequence (seq2seq) voice conversion using non-parallel training data.
D. Erro and A. Moreno, “Frame alignment method for cross-lingual voice conversion,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2007, pp. 1969–1972
1972
Earlier work this paper cites.
D. G. Childers, B. Yegnanarayana, and K. Wu, “Voice conversion: Factors responsible for quality,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 1985, pp. 748–751
1985
Earlier work this paper cites.
L.-J. Liu, Z.-H. Ling, Y. Jiang, M. Zhou, and L.-R. Dai, “WaveNet vocoder with limited training data for voice conversion,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2018, pp. 1983–1987
1987
Earlier work this paper cites.
D. G. Childers, K. Wu, D. M. Hicks, and B. Yegnanarayana, “Voice conversion,” Speech Communication , vol. 8, no. 2, pp. 147–158, 1989
1989
Earlier work this paper cites.
S. Hchreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
A. Kain, “Spectral voice conversion for text-to-speech synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , vol. 1, 1998, pp. 285–288
1998
Earlier work this paper cites.
D. T. Chappell and J. H. L. Hansen, “Speaker-specific pitch contour modeling and modification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , vol. 2, 1998, pp. 885–888
1998
Earlier work this paper cites.
L. M. Arslan, “Speaker transformation algorithm using segmental codebooks (STASC),” Speech Communication , vol. 28, no. 3, pp. 211–226, 1999
1999
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. D. Cheveigné, “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency based F0 extraction: Possible role of a repetitive structure in sounds,” Speech Communication , vol. 27, no. 3–4, pp. 187–207, 1999
1999
Earlier work this paper cites.
J. Kominek and A. W. Black, “CMU ARCTIC databases for speech synthesis,” http://festvox.org/cmu_arctic/index.html
2003
Earlier work this paper cites.
S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in Computer Vision and Pattern Recognition , 2005, pp. 539–546
2005
Earlier work this paper cites.
C.-H. Wu, C.-C. Hsia, T.-H. Liu, and J.-F. Wang, “Voice conversion using duration-embedded bi-HMMs for expressive speech synthesis,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1109–1116, July 2006
2006
Earlier work this paper cites.
H. Duxans, D. Erro, J. Pérez, F. Diego, A. Bonafonte, and A. Moreno, “Voice conversion of non-aligned data using unit selection,” TC-STAR Workshop on Speech-to-Speech Translation , 2006
2006
Earlier work this paper cites.
D. Sundermann, H. Hoge, A. Bonafonte, H. Ney, A. Black, and S. Narayanan, “Text-independent voice conversion based on unit selection,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , vol. 1, 2006, pp. 81–84
2006
Earlier work this paper cites.
M. Müller, “Dynamic time warping,” Information Retrieval for Music and Motion , pp. 69–84, 2007
2007
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,” IEEE Transactions on Audio Speech and Language Processing , vol. 15, no. 8, pp. 2222–2235, 2007
2007
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
S. Desai, E. V. Raghavendra, B. Yegnanarayana, A. W. Black, and K. Prahallad, “Voice conversion using artificial neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2009, pp. 3893–3896
2009
Earlier work this paper cites.
S. Desai, A. W. Black, B. Yegnanarayana, and K. Prahallad, “Spectral mapping using artificial neural networks for voice conversion,” IEEE Transactions on Audio Speech and Language Processing , vol. 18, no. 5, pp. 954–964, 2010
2010
Earlier work this paper cites.
D. Erro, A. Moreno, and A. Bonafonte, “INCA algorithm for training voice conversion systems from nonparallel corpora,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 5, pp. 944–953, 2010
2010
Earlier work this paper cites.
L.-H. Chen, Z.-H. Ling, L.-J. Liu, and L.-R. Dai, “Voice conversion using deep neural networks with layer-wise generative training,” IEEE/ACM Transactions on Audio Speech and Language Processing , vol. 22, no. 12, pp. 1859–1872, 2014
2014
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Advances in Neural Information Processing Systems , 2014, pp. 3104–3112
2014
Cited alongside, same era.
K. Cho, B. Van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder–decoder for statistical machine translation,” in Empirical Methods in Natural Language Processing , 2014, pp. 1724–1734
2014
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Computer Science , 2014
2014
Cited alongside, same era.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from unaligned corpora using variational autoencoding Wasserstein generative adversarial networks,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2017, pp. 3364–3368
2017
Later among the works it cites.
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al. , “Tacotron: Towards end-to-end speech synthesis,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2017, pp. 4006–4010
2017
Later among the works it cites.
C. Veaux, J. Yamagishi, K. MacDonald et al. , “CSTR VCTK corpus: English multi-speaker corpus for cstr voice cloning toolkit,” University of Edinburgh. The Centre for Speech Technology Research (CSTR) , 2017
2017
Later among the works it cites.
T. Kaneko and H. Kameoka, “CycleGAN-VC: Non-parallel voice conversion using cycle-consistent adversarial networks,” in European Signal Processing Conference (EUSIPCO) , 2018, pp. 2114–2117
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Sun, S. Kang, K. Li, and H. Meng, “Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 4869–4873
2015
Cited alongside, same era.
T. Nakashika, T. Takiguchi, and Y. Ariki, “Voice conversion using RNN pre-trained by recurrent temporal restricted Boltzmann machines,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 23, no. 3, pp. 580–587, 2015
2015
Cited alongside, same era.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in International Conference on Learning Representations , 2015
2015
Cited alongside, same era.
T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” in Empirical Methods in Natural Language Processing , 2015, pp. 1412–1421
2015
Cited alongside, same era.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Advances in Neural Information Processing Systems , 2015, pp. 577–585
2015
Cited alongside, same era.
T. Nakashika, T. Takiguchi, Y. Minami, T. Nakashika, T. Takiguchi, and Y. Minami, “Non-parallel training in voice conversion using an adaptive restricted Boltzmann machine,” IEEE/ACM Transactions on Audio, Speech and Language Processing , vol. 24, no. 11, pp. 2032–2045, 2016
2016
Cited alongside, same era.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in 2016 IEEE International Conference on Multimedia and Expo (ICME) , 2016, pp. 1–6
2016
Cited alongside, same era.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from non-parallel corpora using variational auto-encoder,” in 2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA) , 2016, pp. 1–6
2016
Cited alongside, same era.
2018
Later among the works it cites.
F. Fang, J. Yamagishi, I. Echizen, and J. Lorenzo-Trueba, “High-quality nonparallel voice conversion based on cycle-consistent adversarial network,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5279–5283
2018
Later among the works it cites.
S. Liu, J. Zhong, L. Sun, X. Wu, X. Liu, and H. Meng, “Voice conversion across arbitrary speakers based on a single target-speaker utterance,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2018, pp. 496–500
2018
Later among the works it cites.
Y. Saito, Y. Ijima, K. Nishida, and S. Takamichi, “Non-parallel voice conversion using variational autoencoders conditioned by phonetic posteriorgrams and d-vectors,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5274–5278
2018
Later among the works it cites.
J.-c. Chou, C.-c. Yeh, H.-y. Lee, and L.-s. Lee, “Multi-target voice conversion without parallel data by adversarially learning disentangled audio representations,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2018, pp. 501–505
2018
Later among the works it cites.
S. O. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in Advances in Neural Information Processing Systems , 2018, pp. 10 040–10 050
2018
Later among the works it cites.
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Advances in Neural Information Processing Systems , 2018, pp. 4485–4495
2018
Later among the works it cites.
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, “Fitting new speakers based on a short untranscribed sample,” in International Conference on Machine Learning , 2018, pp. 3683–3691
2018
Later among the works it cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerry-Ryan et al. , “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4779–4783
2018
Later among the works it cites.
J.-X. Zhang, Z.-H. Ling, and L.-R. Dai, “Forward attention in sequence-to-sequence acoustic modeling for speech synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4789–4793
2018
Later among the works it cites.
J.-X. Zhang, Z.-H. Ling, L.-J. Liu, Y. Jiang, and L.-R. Dai, “Sequence-to-sequence acoustic modeling for voice conversion,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 3, pp. 631–644, 2019
2019
Closest in time.
K. Tanaka, H. Kameoka, T. Kaneko, and N. Hojo, “ATTS2S-VC: Sequence-to-sequence voice conversion with attention and context preservation mechanisms,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019
2019
Closest in time.
J.-X. Zhang, Z.-H. Ling, Y. Jiang, L.-J. Liu, C. Liang, and L.-R. Dai, “Improving sequence-to-sequence acoustic modeling by adding text-supervision,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6785–6789
2019
Closest in time.
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “CycleGAN-VC2:improved CycleGAN-based non-parallel voice conversion,” in IEEE International Conference on Acoustics Speech and Signal Processing Proceedings , 2019, pp. 6820–6824
2019
Closest in time.
A. Polyak and L. Wolf, “Attention-based WaveNet autoencoder for universal voice conversion,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019
2019
Closest in time.
O. Ocal, O. H. Elibol, G. Keskin, C. Stephenson, A. Thomas, and K. Ramchandran, “Adversarially trained autoencoders for parallel-data-free voice conversion,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019
2019
Closest in time.
H. Zhou, Y. Liu, Z. Liu, P. Luo, and X. Wang, “Talking face generation by adversarially disentangled audio-visual representation,” in AAAI Conference on Artificial Intelligence (AAAI) , 2019
2019
Closest in time.