Fetching the paper…
Reading the bibliography…
Recent singing voice synthesis and conversion advancements necessitate robust singing voice deepfake detection (SVDD) models.
L. Van der Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proc. ICCV , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” Proc. NeurIPS , vol. 32, 2019
2019
Earlier work this paper cites.
P. Lu, J. Wu, J. Luan, X. Tan, and L. Zhou, “XiaoiceSing: A high-quality and integrated singing voice synthesis system,” in Proc. Interspeech , 2020, pp. 1306–1310
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
X. Wang, J. Yamagishi, M. Todisco, H. Delgado, A. Nautsch, N. Evans, M. Sahidullah, V. Vestman, T. Kinnunen, K. A. Lee et al. , “ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech,” Computer Speech & Language , vol. 64, p. 101114, 2020
2020
Earlier work this paper cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in Proc. ICLR , 2020
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Proc. NeurIPS , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
J.-w. Jung, S.-b. Kim, H.-j. Shim, J.-h. Kim, and H.-J. Yu, “Improved rawnet with feature map scaling for text-independent speaker verification using raw waveforms,” Proc. Interspeech , pp. 3583–3587, 2020
2020
Earlier work this paper cites.
I. Ogawa and M. Morise, “Tohoku kiritan singing database: A singing database for statistical parametric singing synthesis using japanese pop songs,” Acoustical Science and Technology , vol. 42, no. 3, pp. 140–145, 2021
2021
Earlier work this paper cites.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in Proc. ICML . PMLR, 2021, pp. 5530–5540
2021
Earlier work this paper cites.
J. Shi, S. Guo, N. Huo, Y. Zhang, and Q. Jin, “Sequence-to-sequence singing voice synthesis with perceptual entropy loss,” in Proc. IEEE ICASSP , 2021, pp. 76–80
2021
Earlier work this paper cites.
S.-W. Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee, “SUPERB: Speech processing universal performance benchmark,” in Proc. Interspeech , 2021, pp. 1194–1198
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
Y. Zhang, J. Cong, H. Xue, L. Xie, P. Zhu, and M. Bi, “VISinger: Variational inference with adversarial learning for end-to-end singing voice synthesis,” in Proc. IEEE ICASSP , 2022, pp. 7237–7241
2022
Cited alongside, same era.
Y. Wang, X. Wang, P. Zhu, J. Wu, H. Li, H. Xue, Y. Zhang, L. Xie, and M. Bi, “Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis,” in Proc. Interspeech , 2022, pp. 4242–4246
Y. Zhang, H. Xue, H. Li, L. Xie, T. Guo, R. Zhang, and C. Gong, “VISinger2: High-fidelity end-to-end singing voice synthesis enhanced by digital signal processing synthesizer,” in Proc. Interspeech , 2023, pp. 4444–4448
2023
Later among the works it cites.
R. Yamamoto, R. Yoneyama, and T. Toda, “NNSVS: A neural network-based singing voice synthesis toolkit,” in Proc. IEEE ICASSP . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
R. Yoneyama, Y.-C. Wu, and T. Toda, “Source-filter HiFi-GAN: Fast and pitch controllable high-fidelity neural vocoder,” in Proc. IEEE ICASSP . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
R. Yamamoto, R. Yoneyama, L. P. Violeta, W.-C. Huang, and T. Toda, “A comparative study of voice conversion models with large-scale speech and singing data: The T13 systems for the singing voice conversion challenge 2023,” in Proc. IEEE ASRU , 2023, pp. 1–6
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
L. Zhang, R. Li, S. Wang, L. Deng, J. Liu, Y. Ren, J. He, R. Huang, J. Zhu, X. Chen, and Z. Zhao, “M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus,” in Proc. NeurIPS (Dataset and Benchmarks Track) , 2022
2022
Cited alongside, same era.
J. Liu, C. Li, Y. Ren, F. Chen, and Z. Zhao, “Diffsinger: Singing voice synthesis via shallow diffusion mechanism,” in Proc. AAAI , vol. 36, no. 10, 2022, pp. 11 020–11 028
2022
Cited alongside, same era.
J. Shi, S. Guo, T. Qian, T. Hayashi, Y. Wu, F. Xu, X. Chang, H. Li, P. Wu, S. Watanabe, and Q. Jin, “Muskits: an end-to-end music processing toolkit for singing voice synthesis,” in Proc. Interspeech , 2022, pp. 4277–4281
2022
Cited alongside, same era.
K. Qian, Y. Zhang, H. Gao, J. Ni, C.-I. Lai, D. Cox, M. Hasegawa-Johnson, and S. Chang, “Contentvec: An improved self-supervised speech representation by disentangling speakers,” in Proc. ICML . PMLR, 2022, pp. 18 003–18 017
2022
Cited alongside, same era.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Cited alongside, same era.
J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. Evans, “AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” in Proc. IEEE ICASSP , 2022, pp. 6367–6371
2022
Cited alongside, same era.
W.-C. Huang, L. P. Violeta, S. Liu, J. Shi, and T. Toda, “The singing voice conversion challenge 2023,” in Proc. IEEE ASRU , 2023, pp. 1–8
2023
Cited alongside, same era.
Timedomain, “ACE Studio.” [Online]. Available: https://acestudio.ai/
Cited in the paper.
T.-h. Feng, A. Dong, C.-F. Yeh, S.-w. Yang, T.-Q. Lin, J. Shi, K.-W. Chang, Z. Huang, H. Wu, X. Chang et al. , “SUPERB@ SLT 2022: Challenge on generalization and efficiency of self-supervised speech representation learning,” in Proc. IEEE SLT , 2023, pp. 1096–1103
2023
Later among the works it cites.
W. Chen, J. Shi, B. Yan, D. Berrebbi, W. Zhang, Y. Peng, X. Chang, S. Maiti, and S. Watanabe, “Joint prediction and denoising for large-scale multilingual self-supervised learning,” in Proc. IEEE ASRU , 2023, pp. 1–8
2023
Later among the works it cites.
2024
Closest in time.
Y. Zang, Y. Zhang, M. Heydari, and Z. Duan, “SingFake: Singing voice deepfake detection,” in Proc. IEEE ICASSP , 2024
2024
Closest in time.
Y. Xie, J. Zhou, X. Lu, Z. Jiang, Y. Yang, H. Cheng, and L. Ye, “FSD: An initial chinese dataset for fake song detection,” in Proc. IEEE ICASSP , 2024
2024
Closest in time.
2024
Closest in time.
J. Shi, H. Inaguma, X. Ma, I. Kulikov, and A. Sun, “Multi-resolution HuBERT: Multi-resolution speech self-supervised learning with masked unit prediction,” in Proc. ICLR , 2024
2024
Closest in time.