Fetching the paper…
Reading the bibliography…
This paper presents a novel framework to build a voice conversion (VC) system by learning from a text-to-speech (TTS) synthesis system, that is called TTS-VC transfer learning.
D. Erro and A. Moreno, “Frame alignment method for cross-lingual voice conversion,” in Proc. INTERSPEECH , 2007, pp. 1969–1972
1972
Earlier work this paper cites.
M. Abe, S. Nakamura, K. Shikano, and H. Kuwabara, “Voice conversion through vector quantization,” Journal of the Acoustical Society of Japan (E) , vol. 11, no. 2, pp. 71–76, 1990
1990
Earlier work this paper cites.
K. Shikano, S. Nakamura, and M. Abe, “Speaker adaptation and voice conversion by codebook mapping,” in IEEE International Sympoisum on Circuits and Systems , 1991, pp. 594–597
1991
Earlier work this paper cites.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in Proc. of IEEE PacRim , vol. 1, 1993, pp. 125–128 vol.1
1993
Earlier work this paper cites.
A. Kain and M. W. Macon, “Spectral voice conversion for text-to-speech synthesis,” in Proc. ICASSP , vol. 1, May 1998, pp. 285–288 vol.1
1998
Earlier work this paper cites.
R. Kuhn, J.-C. Junqua, P. Nguyen, and N. Niedzielski, “Rapid speaker adaptation in eigenvoice space,” IEEE Transactions on Speech and Audio Processing , vol. 8, no. 6, pp. 695–707, 2000
2000
Earlier work this paper cites.
Chung-Hsien Wu, Chi-Chun Hsia, Te-Hsien Liu, and Jhing-Fa Wang, “Voice conversion using duration-embedded bi-HMMs for expressive speech synthesis,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1109–1116, 2006
2006
Earlier work this paper cites.
T. Toda, Y. Ohtani, and K. Shikano, “Eigenvoice conversion based on gaussian mixture model,” in IEEE International Conference on Spoken Language Processing , 2006
2006
Earlier work this paper cites.
O. Turk and L. M. Arslan, “Robust processing techniques for voice conversion,” Computer Speech & Language , vol. 20, no. 4, pp. 441–467, 2006
2006
Earlier work this paper cites.
T. Saitou, M. Goto, M. Unoki, and M. Akagi, “Speech-to-singing synthesis: Converting speaking voices to singing voices by controlling acoustic features unique to singing voices,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics , 2007, pp. 215–218
2007
Earlier work this paper cites.
A. Mouchtaris, Y. Agiomyrgiannakis, and Y. Stylianou, “Conditional vector quantization for voice conversion,” in IEEE ICASSP , vol. 4, 2007, pp. IV–505
2007
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,” IEEE Transactions on Audio, Speech and Language Processing , vol. 15, no. 8, pp. 2222–2235, 2007
2007
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-sne,” Journal of machine learning research , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
T. J. Hazen, W. Shen, and C. White, “Query-by-example spoken term detection using phonetic posteriorgram templates,” in IEEE ASRU , 2009, pp. 421–426
2009
Earlier work this paper cites.
T. L. Nwe, M. Dong, P. Chan, X. Wang, B. Ma, and H. Li, “Voice conversion: From spoken vowels to singing vowels,” in IEEE International Conference on Multimedia and Expo , 2010, pp. 1421–1426
2010
Earlier work this paper cites.
D. Erro, A. Moreno, and A. Bonafonte, “Inca algorithm for training voice conversion systems from nonparallel corpora,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 5, pp. 944–953, 2010
2010
Earlier work this paper cites.
Y. Qian, J. Xu, and F. K. Soong, “A frame mapping based HMM approach to cross-lingual voice transformation,” in IEEE ICASSP , 2011, pp. 5120–5123
2011
Earlier work this paper cites.
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 4, pp. 788–798, May 2011
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The kaldi speech recognition toolkit,” in IEEE ASRU , no. EPFL-CONF-192584, 2011
2011
Earlier work this paper cites.
T. Toda, M. Nakagiri, and K. Shikano, “Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 9, pp. 2505–2517, Nov 2012
2012
Earlier work this paper cites.
K. Nakamura, T. Toda, H. Saruwatari, and K. Shikano, “Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speech,” Speech Communication , vol. 54, no. 1, pp. 134 – 146, 2012. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S0167639311001142
2012
Earlier work this paper cites.
T. Nakashika, R. Takashima, T. Takiguchi, and Y. Ariki, “Voice conversion in high-order eigen space using deep belief nets,” in INTERSPEECH , no. August, pp. 369–372, 2013
2013
Earlier work this paper cites.
X. Tian, Z. Wu, S. W. Lee, and E. S. Chng, “Correlation-based frequency warping for voice conversion,” in Proc. ISCSLP . IEEE, 2014, pp. 211–215
2014
Earlier work this paper cites.
L.-h. Chen, Z.-h. Ling, L.-j. Liu, and L.-r. Dai, “Voice Conversion Using Deep Neural Networks With Layer-Wise Generative Training,” IEEE Transactions on Audio, Speech and Language Processing , vol. 22, no. 12, pp. 1859–1872, 2014
2014
Earlier work this paper cites.
S. H. Mohammadi and A. Kain, “Voice conversion using deep neural networks with speaker-independent pre-training,” in IEEE SLT , 2014, pp. 19–23
2014
Earlier work this paper cites.
Z. Wu, T. Virtanen, E. S. Chng, and H. Li, “Exemplar-based sparse representation with residual compensation for voice conversion,” IEEE/ACM Transactions on Audio, Speech and Language Processing , vol. 22, no. 10, pp. 1506–1521, 2014
2014
Earlier work this paper cites.
2014
Cited alongside, same era.
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in IEEE ICASSP , 2014, pp. 4052–4056
2014
Cited alongside, same era.
B. Sisman, H. Li, and K. C. Tan, “Sparse representation of phonetic features for voice conversion with and without parallel data,” in IEEE ASRU , 2017, pp. 677–684
2014
Cited alongside, same era.
S. Takamichi, T. Toda, A. W. Black, and S. Nakamura, “Modulation spectrum-constrained trajectory training algorithm for gmm-based voice conversion,” in Proc. ICASSP , 2015, pp. 4859–4863
2015
Cited alongside, same era.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, z. Chen, P. Nguyen, R. Pang, I. Lopez Moreno, and Y. Wu, “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds. Curran Associates, Inc., 2018, pp. 4480–4490
2018
Later among the works it cites.
S. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds. Curran Associates, Inc., 2018, pp. 10 019–10 029. [Online]. Available: http://papers.nips.cc/paper/8206-neural-voice-cloning-with-a-few-samples.pdf
2018
Later among the works it cites.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in IEEE ICASSP , 2018, pp. 4879–4883
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Tian, Z. Wu, S. W. Lee, N. Q. Hy, M. Dong, and E. S. Chng, “System fusion for high-performance voice conversion,” Proc. INTERSPEECH , vol. 2015-January, pp. 2759–2763, 2015
2015
Cited alongside, same era.
L. Sun, S. Kang, K. Li, and H. Meng, “Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,” in IEEE ICASSP , 2015, pp. 4869–4873
2015
Cited alongside, same era.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Advances in Neural Information Processing Systems 28 , C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, and R. Garnett, Eds. Curran Associates, Inc., 2015, pp. 577–585. [Online]. Available: http://papers.nips.cc/paper/5847-attention-based-models-for-speech-recognition.pdf
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, F. Bach and D. Blei, Eds., vol. 37. Lille, France: PMLR, 07–09 Jul 2015, pp. 448–456. [Online]. Available: http://proceedings.mlr.press/v37/ioffe15.html
2015
Cited alongside, same era.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from non-parallel corpora using variational auto-encoder,” in IEEE APSIPA , 2016, pp. 1–6
2016
Cited alongside, same era.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in IEEE ICME , 2016, pp. 1–6
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
J.-c. Chou, C.-c. Yeh, H.-y. Lee, and L.-s. Lee, “Multi-target voice conversion without parallel data by adversarially learning disentangled audio representations,” in Proc. INTERSPEECH , 2018, pp. 501–505. [Online]. Available: http://dx.doi.org/10.21437/INTERSPEECH.2018-1830
2018
Later among the works it cites.
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” in Proc. of Machine Learning Research , J. Dy and A. Krause, Eds., vol. 80. Stockholmsmässan, Stockholm Sweden: PMLR, 10–15 Jul 2018, pp. 4693–4702. [Online]. Available: http://proceedings.mlr.press/v80/skerry-ryan18a.html
2018
Later among the works it cites.
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, “The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,” in Proc. Odyssey: The Speaker and Language Recognition Workshop , 2018, pp. 195–202. [Online]. Available: http://dx.doi.org/10.21437/Odyssey.2018-28
2018
Later among the works it cites.
2018
Later among the works it cites.
L.-W. Chen, H.-Y. Lee, and Y. Tsao, “Generative adversarial networks for unpaired voice transformation on impaired speech,” in Proc. INTERSPEECH , 2019, pp. 719–723. [Online]. Available: http://dx.doi.org/10.21437/INTERSPEECH.2019-1265
2019
Later among the works it cites.
B. Sisman, M. Zhang, and H. Li, “Group sparse representation with wavenet vocoder adaptation for spectrum and prosody conversion,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 6, pp. 1085–1097, 2019
2019
Later among the works it cites.
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “Cyclegan-vc2: Improved cyclegan-based non-parallel voice conversion,” in IEEE ICASSP , 2019, pp. 6820–6824
2019
Later among the works it cites.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-shot voice style transfer with only autoencoder loss,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. Long Beach, California, USA: PMLR, 09–15 Jun 2019, pp. 5210–5219. [Online]. Available: http://proceedings.mlr.press/v97/qian19c.html
2019
Later among the works it cites.
H. Luong and J. Yamagishi, “Bootstrapping non-parallel voice conversion from speaker-adaptive text-to-speech,” in IEEE ASRU , 2019, pp. 200–207
2019
Later among the works it cites.
M. Zhang, X. Wang, F. Fang, H. Li, and J. Yamagishi, “Joint training framework for text-to-speech and voice conversion using multi-Source Tacotron and WaveNet,” in Proc. INTERSPEECH , 2019, pp. 1298–1302. [Online]. Available: http://dx.doi.org/10.21437/INTERSPEECH.2019-1357
2019
Later among the works it cites.
2019
Later among the works it cites.
J.-X. Zhang, Z.-H. Ling, Y. Jiang, L.-J. Liu, C. Liang, and L.-R. Dai, “Improving sequence-to-sequence voice conversion by adding text-supervision,” in IEEE ICASSP , 2019, pp. 6785–6789
2019
Later among the works it cites.
J.-X. Zhang, Z.-H. Ling, L.-J. Liu, Y. Jiang, and L.-R. Dai, “Sequence-to-sequence acoustic modeling for voice conversion,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 3, pp. 631–644, 2019
2019
Later among the works it cites.
J. Chorowski, R. Weiss, S. Bengio, and A. Oord, “Unsupervised speech representation learning using wavenet autoencoders,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. PP, pp. 1–1, 09 2019
2019
Later among the works it cites.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,” in Proc. INTERSPEECH , 2019, pp. 1526–1530. [Online]. Available: http://dx.doi.org/10.21437/INTERSPEECH.2019-2441
2019
Later among the works it cites.
M. Zhang, B. Sisman, L. Zhao, and H. Li, “Deepconversion: Voice conversion with limited parallel training data,” Speech Communication , vol. 122, pp. 31 – 43, 2020. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S0167639320302296
2020
Closest in time.
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “CycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-Spectrogram Conversion,” in Proc. Interspeech 2020 , 2020, pp. 2017–2021. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2280
2020
Closest in time.
K. Qian, Z. Jin, M. Hasegawa-Johnson, and G. J. Mysore, “F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6284–6288
2020
Closest in time.
2020
Closest in time.
E. Cooper, C. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, “Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,” in IEEE ICASSP , 2020, pp. 6184–6188
2020
Closest in time.
2020
Closest in time.
J. Zhang, Z. Ling, and L. Dai, “Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 540–552, 2020
2020
Closest in time.