Fetching the paper…
Reading the bibliography…
The increasing prevalence of audio deepfakes poses significant security threats, necessitating robust detection methods.
K. J. Piczak, “ESC: Dataset for Environmental Sound Classification,” in Proceedings of the 23rd Annual ACM Conference on Multimedia (MM) . ACM Press, 2015, pp. 1015–1018
2015
Earlier work this paper cites.
Y. Wang, R. J. Skerry-Ryan et al. , “Tacotron: Towards End-to-End Speech Synthesis,” in INTERSPEECH 2017 , F. Lacerda, Ed. ISCA, 2017, pp. 4006–4010
2017
Earlier work this paper cites.
M. Todisco, H. Delgado, and N. Evans, “Constant Q cepstral coefficients: A spoofing countermeasure for automatic speaker verification,” Computer Speech & Language , vol. 45, pp. 516–535, Sep. 2017
2017
Earlier work this paper cites.
E. A. AlBadawy, S. Lyu, and H. Farid, “Detecting AI-Synthesized Speech Using Bispectral Analysis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , June 2019
2019
Earlier work this paper cites.
G. Lavrentyeva, S. Novoselov et al. , “STC Antispoofing Systems for the ASVspoof2019 Challenge,” in INTERSPEECH 2019 . ISCA, September 2019, pp. 1033–1037
2019
Earlier work this paper cites.
S. Liu, H. Wu et al. , “Adversarial Attacks on Spoofing Countermeasures of Automatic Speaker Verification,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 312–319
2019
Earlier work this paper cites.
X. Wang, J. Yamagishi et al. , “ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech,” Computer Speech & Language , vol. 64, p. 101114, November 2020
2020
Earlier work this paper cites.
Y. Zhang, Z. Jiang et al. , “Black-Box Attacks on Spoofing Countermeasures Using Transferability of Adversarial Examples,” in INTERSPEECH 2020 . ISCA, October 2020, pp. 4238–4242
2020
Earlier work this paper cites.
R. Wang, F. Juefei-Xu et al. , “DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake Voices,” in Proceedings of the 28th ACM International Conference on Multimedia (MM) . Seattle WA USA: ACM, October 2020, pp. 1207–1216
2020
Earlier work this paper cites.
K. He, H. Fan et al. , “Momentum Contrast for Unsupervised Visual Representation Learning,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Seattle, WA, USA: IEEE, June 2020, pp. 9726–9735
2020
Earlier work this paper cites.
T. Chen, S. Kornblith et al. , “A Simple Framework for Contrastive Learning of Visual Representations,” in Proceedings of the 37th International Conference on Machine Learning (ICML) . PMLR, November 2020, pp. 1597–1607
2020
Earlier work this paper cites.
M. Kang and J. Park, “ContraGAN: Contrastive Learning for Conditional Image Generation,” in Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , H. Larochelle, M. Ranzato et al. , Eds., 2020
2020
Earlier work this paper cites.
M. E. Ahmed, I.-Y. Kwak et al. , “Void: A fast and light voice liveness detection system,” in 29th USENIX Security Symposium (USENIX Security 20) , 2020, pp. 2685–2702
2020
Earlier work this paper cites.
J.-w. Jung, S.-b. Kim et al. , “Improved RawNet with Feature Map Scaling for Text-Independent Speaker Verification Using Raw Waveforms,” in INTERSPEECH 2020 . ISCA, Oct. 2020, pp. 1496–1500
2020
Earlier work this paper cites.
E. Wenger, M. Bronckers et al. , “”Hello, It’s Me”: Deep Learning-Based Speech Synthesis Attacks in the Real World,” in Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security (CCS) , New York, NY, USA, 2021, pp. 235–251
2021
Earlier work this paper cites.
H. Tak, J. Patino et al. , “End-to-End anti-spoofing with RawNet2,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6369–6373
2021
Earlier work this paper cites.
G. Hua, A. B. J. Teoh, and H. Zhang, “Towards End-to-End Synthetic Speech Detection,” IEEE Signal Processing Letters , vol. 28, pp. 1265–1269, 2021
2021
Earlier work this paper cites.
J. Yamagishi, X. Wang et al. , “ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection,” in Proc. ASVspoof Challenge workshop , 2021, pp. 47–54
2021
Cited alongside, same era.
A. Saeed, D. Grangier, and N. Zeghidour, “Contrastive Learning of General-Purpose Audio Representations,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Toronto, ON, Canada: IEEE, June 2021, pp. 3875–3879
2021
Cited alongside, same era.
Y. Zhang, F. Jiang, and Z. Duan, “One-Class Learning Towards Synthetic Voice Spoofing Detection,” IEEE Signal Process. Lett. , vol. 28, pp. 937–941, 2021
2021
Cited alongside, same era.
L. Jing and Y. Tian, “Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 43, no. 11, pp. 4037–4058, Nov. 2021
2021
Cited alongside, same era.
E. Conti, D. Salvi et al. , “Deepfake Speech Detection Through Emotion Recognition: A Semantic Approach,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 8962–8966
2022
Later among the works it cites.
S. Park, J. Lee et al. , “Fair Contrastive Learning for Facial Attribute Classification,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 . IEEE, 2022, pp. 10 379–10 388
2022
Later among the works it cites.
L. Blue, K. Warren et al. , “Who Are You (I Really Wanna Know)? Detecting Audio DeepFakes Through Vocal Tract Reconstruction,” in 31st USENIX Security Symposium (USENIX Security 22) , 2022
2022
Later among the works it cites.
Y. A. Li, C. Han et al. , “StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems (NeurIPS) 2023 , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Wang, K. Han et al. , “Contrastive Learning Based Hybrid Networks for Long-Tailed Image Classification,” in IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 . Computer Vision Foundation / IEEE, 2021, pp. 943–952
2021
Cited alongside, same era.
J. Frank and L. Schönherr, “WaveFake: A Data Set to Facilitate Audio Deepfake Detection,” in Proceedings of the Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks 1 , 2021
2021
Cited alongside, same era.
A. K. Singh Yadav, Z. Xiang et al. , “ASSD: Synthetic Speech Detection in the AAC Compressed Domain,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023
2021
Cited alongside, same era.
R. Huang, Z. Zhao et al. , “ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-Speech,” in Proceedings of the 30th ACM International Conference on Multimedia (MM) . Lisboa Portugal: ACM, October 2022, pp. 2595–2605
2022
Cited alongside, same era.
H. Tang, X. Zhang et al. , “Avqvc: One-Shot Voice Conversion By Vector Quantization With Applying Contrastive Learning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2022, pp. 4613–4617
2022
Cited alongside, same era.
Z. Khanjani, G. Watson, and V. P. Janeja, “Audio deepfakes: A survey,” Frontiers Big Data , vol. 5, 2022
2022
Cited alongside, same era.
J.-w. Jung, H.-S. Heo et al. , “AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 6367–6371
2022
Cited alongside, same era.
H. Hojjati and N. Armanfard, “Self-Supervised Acoustic Anomaly Detection Via Contrastive Learning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2022, pp. 3253–3257
2022
Cited alongside, same era.
2023
Later among the works it cites.
Y. A. Li, C. Han, and N. Mesgarani, “Styletts-VC: One-Shot Voice Conversion by Knowledge Transfer From Style-Based TTS Models,” in 2022 IEEE Spoken Language Technology Workshop (SLT) , 2023, pp. 920–927
2023
Later among the works it cites.
J. Li, W. Tu, and L. Xiao, “Freevc: Towards High-Quality Text-Free One-Shot Voice Conversion,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023
2023
Later among the works it cites.
S. Khatsenkova, “Audio deepfake scams: Criminals are using ai to sound like family and people are falling for it,” 2023. [Online]. Available: https://www.euronews.com/embed/2231732
2023
Later among the works it cites.
M. Panariello, W. Ge et al. , “Malafide: a novel adversarial convolutive noise attack against deepfake and spoofing detection systems,” in INTERSPEECH 2023 . ISCA, Aug. 2023, pp. 2868–2872
2023
Later among the works it cites.
J. Guan, F. Xiao et al. , “Anomalous Sound Detection Using Audio Representation with Machine ID Based Contrastive Learning Pretraining,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023
2023
Later among the works it cites.
S. Ding, Y. Zhang, and Z. Duan, “SAMO: Speaker Attractor Multi-Center One-Class Learning For Voice Anti-Spoofing,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, Jun. 2023, pp. 1–5
2023
Later among the works it cites.
J. Zhang, X. Yi, and X. Zhao, “A Compressed Synthetic Speech Detection Method with Compression Feature Embedding,” in INTERSPEECH 2023 . ISCA, Aug. 2023, pp. 5376–5380
2023
Later among the works it cites.
T.-P. Doan, L. Nguyen-Vu et al. , “BTS-E: Audio Deepfake Detection Using Breathing-Talking-Silence Encoder,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, Jun. 2023, pp. 1–5
2023
Later among the works it cites.
X. Liu, F. Zhang et al. , “Self-Supervised Learning: Generative or Contrastive,” IEEE Transactions on Knowledge and Data Engineering , vol. 35, no. 1, pp. 857–876, Jan. 2023
2023
Later among the works it cites.
J. Yi, J. Tao et al. , “ADD 2023: the Second Audio Deepfake Detection Challenge,” in Proceedings of the Workshop on Deepfake Audio Detection and Analysis co-located with 32th International Joint Conference on Artificial Intelligence (IJCAI 2023), Macao, China, August 19, 2023 , vol. 3597, 2023, pp. 125–130
2023
Later among the works it cites.
X. Zhang, J. Yi et al. , “Do You Remember? Overcoming Catastrophic Forgetting for Fake Audio Detection,” in Proceedings of the 40th International Conference on Machine Learning (ICML) . PMLR, Jul. 2023, pp. 41 819–41 831, iSSN: 2640-3498
2023
Later among the works it cites.
——, “What to Remember: Self-Adaptive Continual Learning for Audio Deepfake Detection,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, pp. 19 569–19 577, Mar. 2024
2024
Closest in time.
Z. Ba, Q. Wen et al. , “Transferring Audio Deepfake Detection Capability across Languages,” in Proceedings of the ACM Web Conference (WWW) 2023 . Austin TX USA: ACM, Apr. 2023, pp. 2033–2044
2044
Closest in time.