Fetching the paper…
Reading the bibliography…
The growing prominence of the field of audio deepfake detection is driven by its wide range of applications, notably in protecting the public from potential fraud and other malicious activities, prompting the need for greater attention and research in this area.
A. Bendale and T. E. Boult, “Towards open set deep networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 1563–1572
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 770–778, 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,” IEICE TRANSACTIONS on Information and Systems , vol. 99, no. 7, pp. 1877–1884, 2016
2016
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,” in 2017 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems and Assessment (O-COCOSDA) , 2017, pp. 1–5
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones et al. , “Attention is all you need,” in Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus et al. , Eds., vol. 30. Curran Associates, Inc., 2017
2017
Earlier work this paper cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande et al. , “Efficient neural audio synthesis,” in International Conference on Machine Learning . PMLR, 2018, pp. 2410–2419
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 7132–7141
2018
Earlier work this paper cites.
J.-M. Valin and J. Skoglund, “LPCNet: Improving neural speech synthesis through linear prediction,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5891–5895
2019
Earlier work this paper cites.
J. Hou, Y. Shi, M. Ostendorf, M. Hwang, and L. Xie, “Region proposal network based small-footprint keyword spotting,” IEEE Signal Process. Lett. , vol. 26, no. 10, pp. 1471–1475, 2019. [Online]. Available: https://doi.org/10.1109/LSP.2019.2936282
2019
Earlier work this paper cites.
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe et al. , “Cutmix: Regularization strategy to train strong classifiers with localizable features,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 6023–6032
2019
Earlier work this paper cites.
S.-H. Gao, M.-M. Cheng, K. Zhao, X.-Y. Zhang, M.-H. Yang et al. , “Res2net: A new multi-scale backbone architecture,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 43, no. 2, pp. 652–662, 2019
2019
Earlier work this paper cites.
B. Sisman, J. Yamagishi, S. King, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 132–157, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6199–6203
2020
Earlier work this paper cites.
X. Qin, H. Bu, and M. Li, “Hi-mia: A far-field text-dependent speaker verification database and the baselines,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7609–7613
2020
Earlier work this paper cites.
X. Zhou, Z.-H. Ling, and S. King, “The Blizzard Challenge 2020,” in Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge , vol. 2020, 2020, pp. 1–18
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
Z. Wu, R. K. Das1, J. Yang, and H. Li, “Light convolutional neural network with feature genuinization for detection of synthetic speech attacks,” in Proc. of INTERSPEECH , 2020
2020
Cited alongside, same era.
J. W. Jung, S. B. Kim, H. J. Shim, J. H. Kim, and H. J. Yu, “Improved rawnet with filter-wise rescaling for text-independent speaker verification using raw waveforms,” in Proc. of INTERSPEECH , 2020
2020
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,” in Proc. Interspeech 2020 , 2020, pp. 3830–3834
2020
Cited alongside, same era.
S. Zhao, Q. Yuan, Y. Duan, and Z. Chen, “An end-to-end multi-module audio deepfake generation system for ADD Challenge 2023,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
H. Zhan, Y. Zhang, and X. Yu, “The NeteaseGames system for fake audio generation task of 2023 Audio Deepfake Detection Challenge,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
H. Hua, J. Lu, P. Shi, Z. Shang, Y. Zhang et al. , “Description of a multi-stage audio spoofing system in ADD Challenge 2023,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
Z. Su, J. Liu, Y. Li, Q. Wang, K. Yang et al. , “The Transsion deceptive speech synthesis system for ADD Challenge 2023,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino et al. , “ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection,” in The ASVspoof 2021 Workshop , 2021
2021
Cited alongside, same era.
B. Peng, H. Fan, W. Wang, J. Dong, Y. Li et al. , “DFGC 2021: A deepfake game competition,” in IJCB , 2021
2021
Cited alongside, same era.
A. Mustafa, N. Pia, and G. Fuchs, “StyleMelGAN: An efficient high-fidelity adversarial vocoder with temporal adaptive normalization,” in 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 6034–6038
2021
Cited alongside, same era.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in ICML , 2021
2021
Cited alongside, same era.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-TTS: A diffusion probabilistic model for text-to-speech,” in Proceedings of the 38th International Conference on Machine Learning . PMLR, 2021, pp. 8599–8608
2021
Cited alongside, same era.
J. Yi, Y. Bai, J. Tao, H. Ma, Z. Tian et al. , “Half-truth: A partially fake audio detection dataset,” in Proc. of INTERSPEECH , 2021
2021
Cited alongside, same era.
J. Yi, R. Fu, J. Tao, S. Nie, H. Ma et al. , “ADD 2022: the first audio deep synthesis detection challenge,” in 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022
2022
Cited alongside, same era.
Later among the works it cites.
C. Wu and Y. Wang, “A research on improving the deception ability of speech generated by TTS system,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
H. Wu, Z. Li, L. Xu, Z. Zhang, W. Zhao et al. , “The USTC-NERCSLIP system for the Track 1.2 of Audio Deepfake Detection (ADD 2023) Challenge,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
Y. Zhang, J. Lu, Z. Li, Z. Shang, W. Wang et al. , “Improving the robustness of deepfake audio detection through confidence calibration,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
S. Han, T. Kang, S. Choi, J. Seo, S. Chung et al. , “CAU KU deep fake detection system for ADD 2023 challenge,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
Y. Wang, X. Wang, Y. Chen, Q. Meng, and M. Li, “The DKU-MSXF system description for ADD 2023 Track 1.2,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
Y. Xie, H. Cheng, Y. Wang, and L. Ye, “Single domain generalization for audio deepfake detection,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
Z. Cai, W. Wang, Y. Wang, and M. Li, “The DKU-DUKEECE system for the manipulation region location task of ADD 2023,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
J. Liu, Z. Su, H. Huang, C. Wan, Q. Wang et al. , “TranssionADD: A multi-frame reinforcement based sequence tagging model for audio deepfake detection,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
K. Li, X.-M. Zeng, J.-T. Zhang, and Y. Song, “Convolutional recurrent neural network and multitask learning for manipulation region location,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
J. M. Martín-Doñas and A. Álvarez, “The Vicomtech partial deepfake detection and location system for the 2023 ADD Challenge,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
J. Li, L. Li, M. Luo, X. Wang, S. Qiao et al. , “Multi-grained backend fusion for manipulation region location of partially fake audio,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
J. Lu, Y. Zhang, Z. Li, Z. Shang, W. Wang et al. , “Detecting unknown speech spoofing algorithms with nearest neighbors,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
X. Qin, X. Wang, Y. Chen, Q. Meng, and M. Li, “From speaker verification to deepfake algorithm recognition: Our learned lessons from ADD2023 Track3,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
X.-M. Zeng, J.-T. Zhang, K. Li, Z.-L. Liu, W.-L. Xie et al. , “Deepfake algorithm recognition system with augmented data for ADD 2023 Challenge,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
Z. Wang, Q. Wang, J. Yao, and L. Xie, “The NPU-ASLP system for deepfake algorithm recognition in ADD 2023 Challenge,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.
Y. Tian, Y. Chen, Y. Tang, and B. Fu, “Deepfake algorithm recognition through multi-model fusion based on manifold measure,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis (DADA 2023) , 2023
2023
Later among the works it cites.