Fetching the paper…
Reading the bibliography…
In this paper, we propose a non-parallel any-to-many voice conversion (VC) method termed VoiceGrad.
L.-J. Liu, Z.-H. Ling, Y. Jiang, M. Zhou, and L.-R. Dai, “WaveNet vocoder with limited training data for voice conversion,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2018, pp. 1983–1987
1987
Earlier work this paper cites.
A. Kain and M. W. Macon, “Spectral voice conversion for text-to-speech synthesis,” in Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 1998, pp. 285–288
1998
Earlier work this paper cites.
P. Jax and P. Vary, “Artificial bandwidth extension of speech signals using MMSE estimation based on a hidden Markov model,” in Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2003, pp. 680–683
2003
Earlier work this paper cites.
J. Kominek and A. W. Black, “The CMU Arctic speech databases,” in Proc. ISCA Speech Synthesis Workshop (SSW) , 2004, pp. 223–224
2004
Earlier work this paper cites.
A. Hyvärinen, “Estimation of non-normalized statistical models using score matching,” Journal of Machine Learning Research , vol. 6, pp. 695–709, 2005
2005
Earlier work this paper cites.
A. B. Kain, J.-P. Hosom, X. Niu, J. P. van Santen, M. Fried-Oken, and J. Staehely, “Improving the intelligibility of dysarthric speech,” Speech Communication , vol. 49, no. 9, pp. 743–759, 2007
2007
Earlier work this paper cites.
Z. Inanoglu and S. Young, “Data-driven emotion conversion in spoken English,” Speech Communication , vol. 51, no. 3, pp. 268–283, 2009
2009
Earlier work this paper cites.
D. Felps, H. Bortfeld, and R. Gutierrez-Osuna, “Foreign accent conversion in computer assisted pronunciation training,” Speech Communication , vol. 51, no. 10, pp. 920–932, 2009
2009
Earlier work this paper cites.
O. Türk and M. Schröder, “Evaluation of expressive speech synthesis with voice conversion and copy resynthesis techniques,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 5, pp. 965–973, 2010
2010
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proc. the 13th International Conference on Artificial Intelligence and Statistics , 2010, pp. 249–256
2010
Earlier work this paper cites.
P. Vincent, “A connection between score matching and denoising autoencoders.” Neural Computation , vol. 23, no. 7, pp. 1661–1674, 2011
2011
Earlier work this paper cites.
K. Nakamura, T. Toda, H. Saruwatari, and K. Shikano, “Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speech,” Speech Communication , vol. 54, no. 1, pp. 134–146, 2012
2012
Earlier work this paper cites.
T. Toda, M. Nakagiri, and K. Shikano, “Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 9, pp. 2505–2517, 2012
2012
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” in Proc. International Conference on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
D. P. Kingma, D. J. Rezendey, S. Mohamedy, and M. Welling, “Semi-supervised learning with deep generative models,” in Adv. Neural Information Processing Systems (NIPS) , 2014, pp. 3581–3589
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Adv. Neural Information Processing Systems (NIPS) , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
L. Dinh, D. Krueger, and Y. Bengio, “NICE: Non-linear independent components estimation,” in Proc. International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from non-parallel corpora using variational auto-encoder,” in Proc. Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , 2016, pp. 1–6
2016
Earlier work this paper cites.
H. Zheng, W. Cai, T. Zhou, S. Zhang, and M. Li, “Text-independent voice conversion using deep neural network based phonetic level features,” in Proc. International Conference on Pattern Recognition (ICPR) , 2016, pp. 2872–2877
2016
Earlier work this paper cites.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in 2016 IEEE International Conference on Multimedia and Expo (ICME) , 2016, pp. 1–6
2016
Earlier work this paper cites.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in Advances in Neural Information Processing Systems (NIPS) , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, Eds., vol. 29. Curran Associates, Inc., 2016. [Online]. Available: https://proceedings.neurips.cc/paper/2016/file/ed265bc903a5a097f61d3ec064d96d2e-Paper.pdf
2016
Earlier work this paper cites.
L. Dinh, J. Sohl-Dickstein, and S. Bengio, “Density estimation using real NVP,” in Proc. International Conference on Learning Representations (ICLR) , 2017
2017
Cited alongside, same era.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from unaligned corpora using variational autoencoding Wasserstein generative adversarial networks,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2017, pp. 3364–3368
2017
Cited alongside, same era.
A. van den Oord and O. Vinyals, “Neural discrete representation learning,” in Adv. Neural Information Processing Systems (NIPS) , 2017, pp. 6309–6318
2017
Cited alongside, same era.
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in Proc. International Conference on Computer Vision (ICCV) , 2017, pp. 2223–2232
2017
Cited alongside, same era.
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “StarGAN-VC2: Rethinking conditional methods for stargan-based voice conversion,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2019, pp. 679–683
2019
Later among the works it cites.
2019
Later among the works it cites.
K. Tanaka, H. Kameoka, T. Kaneko, and N. Hojo, “AttS2S-VC: Sequence-to-sequence voice conversion with attention and context preservation mechanisms,” in Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2019, pp. 6805–6809
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Kim, M. Cha, H. Kim, J. K. Lee, and J. Kim, “Learning to discover cross-domain relations with generative adversarial networks,” in Proc. International Conference on Machine Learning (ICML) , 2017, pp. 1857–1865
2017
Cited alongside, same era.
Z. Yi, H. Zhang, P. Tan, and M. Gong, “DualGAN: Unsupervised dual learning for image-to-image translation,” in Proc. International Conference on Computer Vision (ICCV) , 2017, pp. 2849–2857
2017
Cited alongside, same era.
2017
Cited alongside, same era.
H. Miyoshi, Y. Saito, S. Takamichi, and H. Saruwatari, “Voice conversion using sequence-to-sequence learning of context posterior probabilities,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2017, pp. 1268–1272
2017
Cited alongside, same era.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 4835–4839
2017
Cited alongside, same era.
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, “Language modeling with gated convolutional networks,” in Proc. International Conference on Machine Learning (ICML) , 2017, pp. 933–941
2017
Cited alongside, same era.
M. Arjovsky and L. Bottou, “Towards principled methods for training generative adversarial networks,” in Proc. International Conference on Learning Representations (ICLR) , 2017
2017
Cited alongside, same era.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. Courville, “Improved training of Wasserstein GANs,” in Adv. Neural Information Processing Systems (NIPS) , 2017, pp. 5769–5779
2017
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in Adv. Neural Information Processing Systems (NeurIPS) , 2019, pp. 11 918–11 930
2019
Later among the works it cites.
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “Nonparallel voice conversion with augmented classifier star generative adversarial networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2982–2995, 2020
2020
Closest in time.
H. Kameoka, K. Tanaka, D. Kwaśny, and N. Hojo, “ConvS2S-VC: Fully convolutional sequence-to-sequence voice conversion,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 1849–1863, 2020
2020
Closest in time.
2020
Closest in time.
J.-X. Zhang, Z.-H. Ling, and L.-R. Dai, “Recognition-synthesis based non-parallel voice conversion with adversarial learning,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2020, pp. 771–775
2020
Closest in time.
Y. Zhao, W.-C. Huang, X. Tian, J. Yamagishi, R. K. Das, T. Kinnunen, Z. Ling, and T. Toda, “Voice Conversion Challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion,” in Proc. Joint workshop for the Blizzard Challenge and Voice Conversion Challenge , 2020, pp. 80–98
2020
Closest in time.
2020
Closest in time.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems (NeurIPS) , vol. 33, 2020, pp. 6840–6851
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Advances in Neural Information Processing Systems (NeurIPS) , 2020, pp. 17 022–17 033
2020
Closest in time.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 12 449–12 460
2020
Closest in time.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux, “Libri-Light: A benchmark for ASR with limited or no supervision,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7669–7673
2020
Closest in time.
S. Liu, Y. Cao, D. Wang, X. Wu, X. Liu, and H. Meng, “Any-to-many voice conversion with location-relative sequence-to-sequence modeling,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1717–1728, 2021
2021
Closest in time.
A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” in Proceedings of the 38th International Conference on Machine Learning , vol. 139, 2021, pp. 8162–8171
2021
Closest in time.
S. Liu, Y. Cao, D. Su, and H. Meng, “DiffSVC: A diffusion probabilistic model for singing voice conversion,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2021, pp. 741–748
2021
Closest in time.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, M. Kudinov, and J. Wei, “Diffusion-based voice conversion with fast maximum likelihood sampling scheme,” in Proc. International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.
T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-SaruLab system for VoiceMOS Challenge 2022,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , vol. 4521-4525, 2022
2022
Closest in time.
W. C. Huang, E. Cooper, Y. Tsao, H.-M. Wang, T. Toda, and J. Yamagishi, “The VoiceMOS Challenge 2022,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) , 2022, pp. 4536–4540
2022
Closest in time.