Fetching the paper…
Reading the bibliography…
Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021.
“Evaluation of speaker verification security and detection of hmm-based synthetic speech,”
X. Wang, J. Yamagishi, M. Todisco1c, and et al., · 2012
Earlier work this paper cites.
“Asvspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge,”
Z. Wu, T. Kinnunen, N. Evans, J. Yamagishi, C. Hanilc¸i, and et al., · 2015
Earlier work this paper cites.
“A comparison of features for synthetic speech detection,”
M. Sahidullah, T. Kinnunen, and C Hanilçi, · 2015
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, and R. A. Saurous, · 2017
Earlier work this paper cites.
“The asvspoof 2017 challenge: Assessing the limits of replay spoofing attack detection,”
T. Kinnunen, M. Sahidullah, H. Delgado, N. Evans M. Todisco, and et al., · 2017
Earlier work this paper cites.
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
X. Na B. Wu H. Zheng H. Bu, J. Du, · 2017
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
J. Shen, R. Pang, Ron J. Weiss, and et al., · 2018
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Y. Wang, D. Stanton, Y. Zhang, and et al., · 2018
Cited alongside, same era.
“Asvspoof 2019: Future horizons in spoofed and fake audio detection,”
M. Todisco, X. Wang, V. Vestman, Md. Sahidullah, and K. Lee, · 2019
Cited alongside, same era.
“Generalization of audio deepfake detection,”
T. Chen, A. Kumar, P. Nagarsheth, G. Sivaraman, and E. Khoury, · 2020
Cited alongside, same era.
“Deepsonar: Towards effective and robust detection of ai-synthesized fake voices,”
R. Wang, F. Juefei-Xu, Y. Huang, Q. Guo, and et al., · 2020
Cited alongside, same era.
“Aishell-3: A multi-speaker mandarin tts corpus and the baselines,”
Y. Shi, H. Bu, X. Xu, S. Zhang, and M. Li, · 2020
Cited alongside, same era.
“Prosody and voice factorization for few-shot speaker adaptation in the challenge m2voc 2021,”
T. Wang, R. Fu, J. Yi, J. Tao, and S. Wang, · 2021
Later among the works it cites.
“Half-truth: A partially fake audio detection dataset,”
J. Yi, Y. Bai, J. Tao, H. Ma, Z. Tian, C. Wang, T. Wang, and R. Fu, · 2021
Later among the works it cites.
“Continual learning for fake audio detection,”
H. Ma, J. Yi, J. Tao, Y. Bai, Z. Tian, and C. Wang, · 2021
Later among the works it cites.
“Asvspoof 2021: accelerating progress in spoofed and deepfake speech detection,”
J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, X. Liu, K. A. Lee, T. Kinnunen, and N. Evans, · 2021
Later among the works it cites.
“Dfgc 2021: A deepfake game competition,”
B. Peng, H. Fan, W. Wang, J. Dong, Y. Li, S. Lyu, Q. Li, Z. Sun, H. Chen, and B. Chen, · 2021
Later among the works it cites.
“Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Light convolutional neural network with feature genuinization for detection of synthetic speech attacks,”
Z. Wu, R. K. Das1, J. Yang, and H. Li, · 2020
Cited alongside, same era.
“Improved rawnet with filter-wise rescaling for text-independent speaker verification using raw waveforms,”
J. W. Jung, S. B. Kim, H. J. Shim, J. H. Kim, and H. J. Yu, · 2020
Cited alongside, same era.
Y. Fu, L. Cheng, S. Lv, Y. Jv, Y. Kong, Z. Chen, Y. Hu, L. Xie, J. Wu, and H. Bu, · 2021
Later among the works it cites.