Fetching the paper…
Reading the bibliography…
Recent progress in deep generative models has improved the quality of voice conversion in the speech domain.
T. Nakatani, S. Amano, T. Irino, K. Ishizuka, and T. Kondo, “A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,” Speech communication , vol. 50, 2008
2008
Earlier work this paper cites.
F. Villavicencio and J. Bonada, “Applying voice conversion to concatenative singing-voice synthesis,” in Proc. Interspeech , 2010
2010
Earlier work this paper cites.
Z. Duan, H. Fang, B. Li, K. C. Sim, and Y. Wang, “The NUS sung and spoken lyrics corpus: A quantitative comparison of singing and speech,” in Proc. Asia-Pacific Signal and Information Processing Association Annual Summit and Conference , 2013
2013
Earlier work this paper cites.
K. Kobayashi, T. Toda, G. Neubig, S. Sakti, and S. Nakamura, “Statistical singing voice conversion with direct waveform modification based on the spectrum differential,” in Proc. Interspeech , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in Proc. ICLR , 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “World: A vocoder-based high-quality speech synthesis system for real-time applications,” IEICE Transactions on Information and Systems , vol. E99.D, no. 7, pp. 1877–1884, 2016
2016
Earlier work this paper cites.
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio, “SampleRNN: An Unconditional End-to-End Neural Audio Generation Model,” in Proc. ICLR , 2017
2017
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanab, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in Proc. ICASSP , 2017
2017
Earlier work this paper cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in Proc. ICML , 2018
2018
Earlier work this paper cites.
A. Liutkus, F.-R. Stöter, and N. Ito, “The 2018 signal separation evaluation campaign,” in Proc LVA/ICA , 2018
2018
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-End Speech Processing Toolkit,” in Proc. Interspeech , 2018, pp. 2207–2211
2018
Earlier work this paper cites.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in Proc. IEEE PACRIM , 1993
2018
Earlier work this paper cites.
E. Nachmani and L. Wolf, “Unsupervised singing voice conversion,” in Proc.Interspeech , 2019
2019
Earlier work this paper cites.
J.-c. Chou, C.-c. Yeh, and H.-y. Lee, “One-shot voice conversion by separating speaker and content representations with instance normalization,” in Proc. Interspeech , 2019
2019
Earlier work this paper cites.
R. Prenger, R. Valle, and B. Catanzaro, “WaveGlow: A flowbased generative network for speech synthesis,” in Proc. ICASSP , 2019
2019
Earlier work this paper cites.
C. Donahue, J. McAuley, and M. Puckette, “Adversarial audio synthesis,” in Proc. ICLR , 2019
2019
Earlier work this paper cites.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brebisson, Y. Bengio, and A. Courville, “MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesi,” in Proc. NeurIPS , 2019
2019
Earlier work this paper cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in Proc. NeurIPS , 2019
2019
Earlier work this paper cites.
Y. Yazıcı, C.-S. Foo, S. Winkler, K.-H. Yap, G. Piliouras, and V. Chandrasekhar, “The unusual effectiveness of averaging in GAN training,” in Proc. ICLR , 2019
2019
Cited alongside, same era.
C. Deng, C. Yu, H. Lu, C. Weng, and D. Yu, “Pitchnet: Unsupervised singing voice conversion with pitch adversarial network,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
A. Polyak, L. Wolf, Y. Adi, and Y. Taigman, “Unsupervised cross-domain singing voice conversion,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
Y.-J. Luo, C.-C. Hsu, K. Agres, and D. Herremans, “Singing voice conversion with disentangled representations of singer and vocal technique using variational autoencoders,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
L. Zhang, C. Yu, H. Lu, C. Weng, C. Zhang, Y. Wu, X. Xie, Z. Li, and D. Yu, “DurIAN-SC: Duration Informed Attention Network based Singing Voice Conversion System,” in Proc.Interspeech , 2020
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “WaveGrad: Estimating gradients for waveform generation,” in Proc. ICLR , 2021
2021
Later among the works it cites.
Y. A. Li, A. Zare, and N. Mesgarani, “StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
N. Takahashi and Y. Mitsufuji, “Densely connected multidilated convolutional networks for dense prediction tasks,” in Proc. CVPR , 2021
2021
Later among the works it cites.
B. Sharma, X. Gao, K. Vijayan, X. Tian, and H. Li, “NHSS: A Speech and Singing Parallel Database,” Speech Communication , vol. 13, 2021
2021
Later among the works it cites.
C. Wang, Z. Li, B. Tang, X. Yin, Y. Wan, Y. Yu, and Z. Ma, “Towards High-fidelity Singing Voice Conversion with Acoustic Reference and Contrastive Predictive Coding,” in Proc. ACM Multimedia , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
D.-Y. Wu, Y.-H. Chen, and H.-Y. Lee, “Vqvc+: One-shot voice conversion by vector quantization and u-net architecture,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
X. Miao, M. Sun, X. Zhang, and Y. Wang, “Noise-robust voice conversion using high-quefrency boosting via sub-band cepstrum conversion and fusion,” Applied Sciences , vol. 10, no. 1, 2020
2020
Cited alongside, same era.
W. Ping, K. Peng, K. Zhao, and Z. Song, “WaveFlow: A compact flow-based model for raw audio,” in Proc. ICML , 2020
2020
Cited alongside, same era.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. NeurIPS , 2020
2020
Cited alongside, same era.
P. A. Jonathan Ho, Ajay Jain, “Denoising diffusion probabilistic models,” in Proc. NeurIPS , 2020
2020
Cited alongside, same era.
Y. Choi, Y. Uh, J. Yoo, and J.-W. Ha, “StarGAN v2: Diverse Image Synthesis for Multiple Domains,” in Proc. CVPR , 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,” in Proc. ICLR , 2021
2021
Later among the works it cites.
H. Guo, Z. Zhou, F. Meng, and K. Liu, “Improving adversarial waveform generation based singing voice conversion with harmonic signals,” in Proc. ICASSP , 2022
2022
Closest in time.
Y. Zhang, P. Yang, J. Xiao, Y. Bai, H. Che, and X. Wang, “K-converter: An unsupervised singing voice conversion system,” in Proc. ICASSP , 2022
2022
Closest in time.
Y. Zhou and X. Lu, “Hifi-svc: Fast high fidelity cross-domain singing voice conversion,” in Proc. ICASSP , 2022
2022
Closest in time.
X. Li, S. Liu, and Y. Shan, “A Hierarchical Speaker Representation Framework for One-shot Singing Voice Conversion,” in Proc. Interspeech , 2022
2022
Closest in time.
C. Xie, Y.-C. Wu, P. L. Tobing, W.-C. Huang, and T. Toda, “Direct Noisy Speech Modeling for Noisy-To-Noisy Voice Conversion,” in Proc. ICASSP , 2022
2022
Closest in time.
Y.-J. Lu, Z.-Q. Wang, S. Watanabe, A. Richard, C. Yu, and Y. Tsao, “Conditional Diffusion Probabilistic Model for Speech Enhancement,” in Proc. ICASSP , 2022
2022
Closest in time.
S. gil Lee, H. Kim, C. Shin, X. Tan, C. Liu, Q. Meng, T. Qin, W. Chen, S. Yoon, and T.-Y. Liu, “ ”PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior”,” in Proc. ICLR , 2022
2022
Closest in time.
Y. Koizumi, H. Zen, K. Yatabe, N. Chen, and M. Bacchiani, “SpecGrad: Diffusion Probabilistic Model based Neural Vocoder with Adaptive Noise Spectral Shaping,” in Proc. Interspeech , 2022
2022
Closest in time.
J. Liu, C. Li, Y. Ren, F. Chen, and Z. Zhao, “DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism,” in Proc. AAAI , 2022
2022
Closest in time.
J. S. Um, Y. Choi, and H. Kim, “Acnn-vc: Utilizing adaptive convolution neural network for one-shot voice conversion,” in Proc. Interspeech , 2022
2022
Closest in time.
G. Mittag and S. Moller, “Deep Learning Based Assessment of Synthetic Speech Naturalness,” in Proc. Interspeech , 2022
2022
Closest in time.
T. Jayashankar, J. Wu, L. Sari, D. Kant, V. Manohar, and Q. He, “Self-supervised representations for singing voice conversion,” in Proc. ICASSP , 2023
2023
Closest in time.
N. Takahashi, M. K. Singh, and Y. Mitsufuji, “Hierarchical Diffusion Models for Singing Voice Neural Vocoder,” in Proc. ICASSP , 2023
2023
Closest in time.