Fetching the paper…
Reading the bibliography…
Audio deepfake detection is an emerging active topic.
D. A. Reynolds and R. C. Rose, “Robust text-independent speaker identification using gaussian mixture speaker models,” IEEE Trans. Speech Audio Process. , vol. 3, pp. 72–83, 1995
1995
Earlier work this paper cites.
A. J. Hunt and A. W. Black, “Unit selection in a concatenative speech synthesis system using a large speech database,” in IEEE International Conference on Acoustics , 1996
1996
Earlier work this paper cites.
L. Rabiner and B.-H. Juang, Fundamentals of speech recognition . Fundamentals of speech recognition, 1999
1999
Earlier work this paper cites.
Y. Stylianou, “Applying the harmonic plus noise model in concatenative speech synthesis,” Speech & Audio Processing IEEE Transactions on , vol. 9, no. 1, pp. 21–29, 2001
2001
Earlier work this paper cites.
T. Ojala, M. Pietikäinen, and T. Mäenpää, “Multiresolution gray-scale and rotation invariant texture classification with local binary patterns,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 24, pp. 971–987, 2002
2002
Earlier work this paper cites.
G. Zhao and M. Pietikäinen, “Dynamic texture recognition using local binary patterns with an application to facial expressions,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 29, pp. 915–928, 2007
2007
Earlier work this paper cites.
S. Chakroborty, A. Roy, and G. Saha, “Improved closed set text-independent speaker identification by combining mfcc with evidence from flipped filter banks,” World Academy of Science, Engineering and Technology, International Journal of Electrical, Computer, Energetic, Electronic and Communication Engineering , vol. 2, pp. 2554–2561, 2008
2008
Earlier work this paper cites.
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE Transactions on Neural Networks , vol. 20, pp. 61–80, 2009
2009
Earlier work this paper cites.
P. L. D. Leon, M. Pucher, and J. Yamagishi, “Evaluation of the vulnerability of speaker verification to synthetic speech,” in Proc. of Odyssey: The Speaker and Language Recognition Workshop , 2010
2010
Earlier work this paper cites.
L. Chen, W. Guo, and L. Dai, “Speaker verification against synthetic speech,” 2010 7th International Symposium on Chinese Spoken Language Processing , pp. 309–312, 2010
2010
Earlier work this paper cites.
T. K. A and H. L. B, “An overview of text-independent speaker recognition: From features to supervectors,” Speech Communication , vol. 52, no. 1, pp. 12–40, 2010
2010
Earlier work this paper cites.
E. Godoy, O. Rosec, and T. Chonavel, “Voice conversion using dynamic frequency warping with amplitude scaling, for parallel or nonparallel corpora,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 4, pp. 1313–1323, 2011
2011
Earlier work this paper cites.
C.-C. Chang and C.-J. Lin, “Libsvm: A library for support vector machines,” ACM Trans. Intell. Syst. Technol. , vol. 2, pp. 27:1–27:27, 2011
2011
Earlier work this paper cites.
E. S. C. Zhizheng Wu and H. Li, “Detecting converted speech and natural speech for anti-spoofing attack in speaker recognition,” in Interspeech , 2012
2012
Earlier work this paper cites.
P. Leon, B. Stewart, and J. Yamagishi, “Synthetic speech discrimination using pitch pattern statistics derived from image analysis,” in Interspeech , 2012
2012
Earlier work this paper cites.
F. Alegre, R. Vipperla, and N. W. D. Evans, “Spoofing countermeasures for the protection of automatic speaker recognition systems against attacks with artificial signals,” in Interspeech , 2012
2012
Earlier work this paper cites.
S. Prince, P. Li, Y. Fu, U. Mohammed, and J. H. Elder, “Probabilistic models for inference about identity,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 34, pp. 144–157, 2012
2012
Earlier work this paper cites.
Z. Wu, C. E. Siong, and H. Li, “Detecting converted speech and natural speech for anti-spoofing attack in speaker recognition,” in Interspeech , 2012
2012
Earlier work this paper cites.
H. Zen, A. W. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” 2013 IEEE International Conference on Acoustics, Speech and Signal Processing , pp. 7962–7966, 2013
2013
Earlier work this paper cites.
F. Alegre, A. Amehraye, and N. Evans, “A one-class classification approach to generalised speaker verification spoofing countermeasures using local binary patterns,” in Proc. of Int. Conf. on Biometrics: Theory, Applications and Systems (BTAS) , 2013
2013
Earlier work this paper cites.
Z. Kons and H. Aronowitz, “Voice transformation-based spoofing of text dependent speaker verification systems,” in Annual Conference of the International Speech Communication Association (Interspeech) , 2013
2013
Earlier work this paper cites.
Z. Wu, A. Larcher, K. A. Lee, and et al., “Vulnerability evaluation of speaker verification under voice conversion spoofing: the effect of text constraints,” in Annual Conference of the International Speech Communication Association (Interspeech) , 2013
2013
Earlier work this paper cites.
Z. Wu, X. Xiong, E. S. Chng, and H. Li, “Synthetic speech detection using temporal modulation feature,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2013
2013
Earlier work this paper cites.
F. Alegre, R. Vipperla, A. Amehraye, and N. W. D. Evans, “A new speaker verification spoofing countermeasure based on local binary patterns,” in Interspeech , 2013
2013
Earlier work this paper cites.
F. Alegre, A. Amehraye, and N. W. D. Evans, “A one-class classification approach to generalised speaker verification spoofing countermeasures using local binary patterns,” 2013 IEEE Sixth International Conference on Biometrics: Theory, Applications and Systems (BTAS) , pp. 1–8, 2013
2013
Earlier work this paper cites.
T. N. Sainath, B. Kingsbury, A. rahman Mohamed, and B. Ramabhadran, “Learning filter banks within a deep neural network framework,” 2013 IEEE Workshop on Automatic Speech Recognition and Understanding , pp. 297–302, 2013
2013
Earlier work this paper cites.
J. Andén and S. Mallat, “Deep scattering spectrum,” IEEE Transactions on Signal Processing , vol. 62, pp. 4114–4128, 2013
2013
Earlier work this paper cites.
T. B. Amin, J. S. German, and P. Marziliano, “Detecting voice disguise from speech variability: Analysis of three glottal and vocal tract measures,” Journal of the Acoustical Society of America , vol. 134, pp. 4068–4068, 2013
2013
Earlier work this paper cites.
K. Kasi and S. A. Zahorian, “Yet another algorithm for pitch tracking,” in 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing , 2002, pp. I–361–I–364
2014
Earlier work this paper cites.
Z. Wu, N. Evans, T. Kinnunen, J. Yamagishi, F. Alegre, and H. Li, “Spoofing and countermeasures for speaker verification: A survey,” Speech Communication , vol. 66, pp. 130–153, 2015
2015
Earlier work this paper cites.
Z. Wu, T. Kinnunen, N. Evans, J. Yamagishi, C. Hanilc¸i, and et al., “Asvspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge,” in Proc. of INTERSPEECH , 2015
2015
Earlier work this paper cites.
Z. Wu, A. Khodabakhsh, C. Demiroglu, and et al., “Sas : A speaker verification spoofing database containing diverse attacks,” in IEEE International Conference on Acoustics, Speech and Signal , 2015
2015
Earlier work this paper cites.
M. Sahidullah, T. Kinnunen, and C. Hanilçi, “A comparison of features for synthetic speech detection,” in Proc. of INTERSPEECH , 2015
2015
Earlier work this paper cites.
X. Xiao, X. Tian, S. Du, H. Xu, and H. Li, “Spoofing speech detection using high dimensional magnitude and phase features: the ntu approach for asvspoof 2015 challenge.” in Interspeech , 2015
2015
Earlier work this paper cites.
J. Sanchez, I. Saratxaga, I. Hernaez, E. Navas, D. Erro, and T. Raitio, “Toward a universal synthetic speech spoofing detection using phase information,” IEEE Transactions on Information Forensics & Security , vol. 10, no. 4, pp. 810–820, 2015
2015
Earlier work this paper cites.
T. B. Patel and H. Patil, “Combining evidences from mel cepstral, cochlear filter cepstral and instantaneous frequency features for detection of natural vs. spoofed speech,” in Conference of International Speech Communication Association , 2015
2015
Earlier work this paper cites.
N. Chen, Y. Qian, H. Dinkel, B. Chen, and K. Yu, “Robust deep feature for spoofing detection - the sjtu system for asvspoof 2015 challenge,” in Interspeech , 2015
2015
Earlier work this paper cites.
J. A. V. López, A. Miguel, A. Ortega, and E. L. SOLANO, “Spoofing detection with dnn and one-class svm for the asvspoof 2015 challenge,” in Interspeech , 2015
2015
Earlier work this paper cites.
A. Sizov, E. el Khoury, T. H. Kinnunen, Z. Wu, and S. Marcel, “Joint speaker verification and antispoofing in the $i$ -vector space,” IEEE Transactions on Information Forensics and Security , vol. 10, pp. 821–832, 2015
2015
Earlier work this paper cites.
C. Hanilçi, T. H. Kinnunen, M. Sahidullah, and A. Sizov, “Classifiers for synthetic speech detection: a comparison,” in Interspeech , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 770–778, 2015
2015
Earlier work this paper cites.
Z. Jin, A. Finkelstein, S. DiVerdi, J. Lu, and G. J. Mysore, “Cute: A concatenative method for voice conversion using exemplar-based unit selection,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 5660–5664
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
X. Tian, Z. Wu, X. Xiong, E. S. Chng, and H. Li, “Spoofing detection from a feature representation perspective,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016
2016
Earlier work this paper cites.
S. Novoselov, A. Kozlov, G. Lavrentyeva, K. Simonchik, and V. Shchemelinin, “Stc anti-spoofing systems for the asvspoof 2015 challenge,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016
2016
Earlier work this paper cites.
Z. Wu, P. L. De Leon, C. Demiroglu, A. Khodabakhsh, S. King, Z.-H. Ling, D. Saito, B. Stewart, T. Toda, M. Wester, and J. Yamagishi, “Anti-spoofing for text-independent speaker verification: An initial database, comparison of countermeasures, and human performance,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 4, pp. 768–783, 2016
2016
Earlier work this paper cites.
M. Todisco, H. Delgado, and N. Evans, “A new feature for automatic speaker verification antispoofing: Constant q cepstral coefficients,” in Processings of Odyssey 2016 , 2016
2016
Earlier work this paper cites.
Y. Qian, N. Chen, and K. Yu, “Deep features for automatic spoofing detection,” Speech Commun. , vol. 85, pp. 43–52, 2016
2016
Earlier work this paper cites.
P. Korshunov and S. Marcel, “Cross-database evaluation of audio-based spoofing detection systems,” in Interspeech , 2016
2016
Earlier work this paper cites.
T. Kinnunen, M. Sahidullah, H. Delgado, N. E. M. Todisco, and et al., “The asvspoof 2017 challenge: Assessing the limits of replay spoofing attack detection,” in Proc. of INTERSPEECH , 2017
2017
Earlier work this paper cites.
T. Kinnunen, M. Sahidullah, H. Delgado, N. E. M. Todisco, and et al., “The asvspoof 2017 challenge: Assessing the limits of replay spoofing attack detection,” in Annual Conference of the International Speech Communication Association (Interspeech) , 2017
2017
Earlier work this paper cites.
G. Lavrentyeva, S. Novoselov, E. Malykh, A. Kozlov, and V. Shchemelinin, “Audio replay attack detection with deep learning frameworks,” in Interspeech 2017 , 2017
2017
Earlier work this paper cites.
Y. Hong, Z. H. Tan, Z. Ma, and J. Guo, “Dnn filter bank cepstral coefficients for spoofing detection,” IEEE Access , vol. 5, no. 99, pp. 4779–4787, 2017
2017
Earlier work this paper cites.
H. B. Sailor, D. M. Agrawal, and H. A. Patil, “Unsupervised filterbank learning using convolutional restricted boltzmann machine for environmental sound classification,” in Interspeech , 2017
2017
Earlier work this paper cites.
Z. Ji, Z.-Y. Li, P. Li, M. An, S. Gao, D. Wu, and F. Zhao, “Ensemble learning for countermeasure of audio replay spoofing attack in asvspoof2017,” in Interspeech , 2017
2017
Earlier work this paper cites.
M. Todisco, H. Delgado, and N. Evans, “Constant q cepstral coefficients: A spoofing countermeasure for automatic speaker verification,” Computer Speech and Language , vol. 45, pp. 516–535, 2017
2017
Earlier work this paper cites.
J. Hu, L. Shen, S. Albanie, G. Sun, and E. Wu, “Squeeze-and-excitation networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 42, pp. 2011–2023, 2017
2017
Earlier work this paper cites.
H. Muckenhirn, M. Magimai.-Doss, and S. Marcel, “End-to-end convolutional neural network-based voice presentation attack detection,” 2017 IEEE International Joint Conference on Biometrics (IJCB) , pp. 335–341, 2017
2017
Cited alongside, same era.
H. Dinkel, N. Chen, Y. Qian, and K. Yu, “End-to-end spoofing detection with raw waveform cldnns,” 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 4860–4864, 2017
2017
Cited alongside, same era.
M. Todisco, H. Delgado, K.-A. Lee, M. Sahidullah, N. W. D. Evans, T. H. Kinnunen, and J. Yamagishi, “Integrated presentation attack detection and automatic speaker verification: Common features and gaussian back-end fusion,” in Interspeech , 2018
2018
Cited alongside, same era.
M. Pal, D. Paul, and G. Saha, “Synthetic speech detection using fundamental frequency variation and spectral features,” Computer Speech & Language , vol. 48, pp. 31–50, 2018
2018
Cited alongside, same era.
2021
Later among the works it cites.
Z. Zhang, Y. Gu, X. Yi, and X. Zhao, “Fmfcc-a: A challenging mandarin dataset for synthetic speech detection,” in International Workshop on Digital Watermarking , 2021
2021
Later among the works it cites.
X. Wang and J. Yamagishi, “Investigating self-supervised front ends for speech spoofing countermeasures,” in The Speaker and Language Recognition Workshop , 2021
2021
Later among the works it cites.
N. Zeghidour, O. Teboul, F. Quitry, and M. Tagliasacchi, “Leaf: A learnable frontend for audio classification,” in ICLR , 2021
2021
Later among the works it cites.
Y. Gao, T. Vuong, M. Elyasi, G. Bharaj, and R. Singh, “Generalized spoofing detection inspired from audio generation artifacts,” 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Zeghidour, N. Usunier, I. Kokkinos, T. Schatz, G. Synnaeve, and E. Dupoux, “Learning filterbanks from raw speech for phone recognition,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 5509–5513, 2018
2018
Cited alongside, same era.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with sincnet,” 2018 IEEE Spoken Language Technology Workshop (SLT) , pp. 1021–1028, 2018
2018
Cited alongside, same era.
C.-I. Lai, A. Abad, K. Richmond, J. Yamagishi, N. Dehak, and S. King, “Attentive filtering networks for audio replay attack detection,” ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6316–6320, 2018
2018
Cited alongside, same era.
X. Wu, R. He, Z. Sun, and T. Tan, “A light cnn for deep face representation with noisy labels,” IEEE Transactions on Information Forensics and Security , vol. 13, no. 11, pp. 2884–2896, 2018
2018
Cited alongside, same era.
B. Bitesize, “Deepfakes: What are they and why would i make one?” 2019. [Online]. Available: https://www.bbc.co.uk/bitesize/articles/zfkwcqt
2019
Cited alongside, same era.
C. Stupp, “Fraudsters used ai to mimic ceo’s voice in unusual cybercrime case,” 2019. [Online]. Available: https://www.wsj.com/articles/fraudsters-use-ai-to-mimic-ceos-voice-in-unusual-cybercrime-case-11567157402
2019
Cited alongside, same era.
S. Dahmani, V. Colotte, V. Girard, and S. Ouni, “Conditional variational auto-encoder for text-driven expressive audiovisual speech synthesis,” in INTERSPEECH 2019-20th Annual Conference of the International Speech Communication Association , 2019
2019
Cited alongside, same era.
M. Todisco, X. Wang, V. Vestman, M. Sahidullah, and K. Lee, “Asvspoof 2019: Future horizons in spoofed and fake audio detection,” in Proc. of INTERSPEECH , 2019
2019
Cited alongside, same era.
2021
Later among the works it cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Later among the works it cites.
Q. Fu, Z. Teng, J. White, M. G. Powell, and D. C. Schmidt, “Fastaudio: A learnable audio front-end for spoof speech detection,” ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 3693–3697, 2021
2021
Later among the works it cites.
H. Tak, J. Patino, M. Todisco, A. Nautsch, N. W. D. Evans, and A. Larcher, “End-to-end anti-spoofing with rawnet2,” ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6369–6373, 2021
2021
Later among the works it cites.
A. Tomilov, A. F. Svishchev, M. Volkova, A. Chirkovskiy, A. S. Kondratev, and G. Lavrentyeva, “Stc antispoofing systems for the asvspoof2021 challenge,” 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge , 2021
2021
Later among the works it cites.
X. Wang and J. Yamagishi, “Investigating self-supervised front ends for speech spoofing countermeasures,” in The Speaker and Language Recognition Workshop , 2021
2021
Later among the works it cites.
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. M. Pino, A. Baevski, A. Conneau, and M. Auli, “Xls-r: Self-supervised cross-lingual speech representation learning at scale,” in Interspeech , 2021
2021
Later among the works it cites.
Y. Xie, Z. Zhang, and Y. Yang, “Siamese network with wav2vec feature for spoofing speech detection,” in Interspeech , 2021
2021
Later among the works it cites.
R. Hemavathi and R. Kumaraswamy, “Voice conversion spoofing detection by exploring artifacts estimates,” Multimedia Tools and Applications , vol. 80, pp. 23 561 – 23 580, 2021
2021
Later among the works it cites.
T. Chen, E. el Khoury, K. Phatak, and G. Sivaraman, “Pindrop labs’ submission to the asvspoof 2021 challenge,” 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge , 2021
2021
Later among the works it cites.
X. Li, N. Li, C. Weng, X. Liu, D. Su, D. Yu, and H. M. Meng, “Replay and synthetic speech detection with res2net architecture,” ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6354–6358, 2021
2021
Later among the works it cites.
Y. Zhang, W. Wang, and P. Zhang, “The effect of silence and dual-band fusion in anti-spoofing system,” in Interspeech , 2021
2021
Later among the works it cites.
H. Tak, J. weon Jung, J. Patino, M. Todisco, and N. W. D. Evans, “Graph attention networks for anti-spoofing,” in Interspeech , 2021
2021
Later among the works it cites.
W. Ge, M. Panariello, J. Patino, M. Todisco, and N. W. D. Evans, “Partially-connected differentiable architecture search for deepfake and spoofing detection,” in Interspeech , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
W. Ge, J. Patino, M. Todisco, and N. W. D. Evans, “Raw differentiable architecture search for speech deepfake and spoofing detection,” 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge , 2021
2021
Later among the works it cites.
A. K. Singh and P. Singh, “Detection of ai-synthesized speech using cepstral & bispectral statistics,” 2021 IEEE 4th International Conference on Multimedia Information Processing and Retrieval (MIPR) , pp. 412–417, 2021
2021
Later among the works it cites.
I.-Y. Kwak, S. Kwag, J. Lee, J. H. Huh, C.-H. Lee, Y. B. Jeon, J.-H. Hwang, and J. W. Yoon, “Resmax: Detecting voice spoofing attacks with residual network and max feature map,” 2020 25th International Conference on Pattern Recognition (ICPR) , pp. 4837–4844, 2021
2021
Later among the works it cites.
X. Wang and J. Yamagishi, “A comparative study on recent neural spoofing countermeasures for synthetic speech detection,” in Interspeech , 2021
2021
Later among the works it cites.
G. Hua, A. Teoh, and H. Zhang, “Towards end-to-end synthetic speech detection,” IEEE Signal Processing Letters , vol. 28, pp. 1265–1269, 2021
2021
Later among the works it cites.
Y. Ma, Z. Ren, and S. Xu, “Rw-resnet: A novel speech anti-spoofing model using raw waveform,” Interspeech , 2021
2021
Later among the works it cites.
A. Nautsch, X. Wang, N. W. D. Evans, T. H. Kinnunen, V. Vestman, M. Todisco, H. Delgado, M. Sahidullah, J. Yamagishi, and K.-A. Lee, “Asvspoof 2019: Spoofing countermeasures for the detection of synthesized, converted and replayed speech,” IEEE Transactions on Biometrics, Behavior, and Identity Science , vol. 3, pp. 252–265, 2021
2021
Later among the works it cites.
H. Ma, J. Yi, J. Tao, Y. Bai, Z. Tian, and C. Wang, “Continual learning for fake audio detection,” in Proc. of INTERSPEECH , 2021
2021
Later among the works it cites.
Z. Almutairi and H. Elgibreen, “A review of modern audio deepfake detection methods: Challenges and future directions,” Algorithms , 2022
2022
Later among the works it cites.
R. Huang, M. W. Y. Lam, J. Wang, D. Su, D. Yu, Y. Ren, and Z. Zhao, “Fastdiff: A fast conditional diffusion model for high-quality speech synthesis,” in International Joint Conference on Artificial Intelligence , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Yi, R. Fu, J. Tao, S. Nie, H. Ma, C. Wang, T. Wang, Z. Tian, Y. Bai, C. Fan, S. Liang, S. Wang, S. Zhang, X. Yan, L. Xu, Z. Wen, and H. Li, “Add 2022: the first audio deep synthesis detection challenge,” in 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Xue, C. Fan, Z. Lv, J. Tao, J. Yi, C. Zheng, Z. Wen, M. Yuan, and S. Shao, “Audio deepfake detection based on a combination of F0 information and real plus imaginary spectrogram features,” in DDAM@MM 2022: Proceedings of the 1st International Workshop on Deepfake Detection for Audio Multimedia, Lisboa, Portugal, 14 October 2022 . ACM, 2022, pp. 19–26
2022
Later among the works it cites.
E. Conti, D. Salvi, C. Borrelli, B. Hosler, P. Bestagini, F. Antonacci, A. Sarti, M. C. Stamm, and S. Tubaro, “Deepfake speech detection through emotion recognition: A semantic approach,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Singapore, 23-27 May 2022 . IEEE, 2022, pp. 8962–8966
2022
Later among the works it cites.
J.-Y. Pan, S. Nie, H. Zhang, S. He, K. Zhang, S. Liang, X. Zhang, and J. Tao, “Speaker recognition-assisted robust audio deepfake detection,” in INTERSPEECH , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
J. M. Mart’in-Donas and A. Álvarez, “The vicomtech audio deepfake detection system based on wav2vec2 for the 2022 add challenge,” 2022, pp. 9241–9245
2022
Later among the works it cites.
Z. Lv, S. Zhang, K. Tang, and P. Hu, “Fake audio detection based on unsupervised pretraining models,” 2022, pp. 9231–9235
2022
Later among the works it cites.
J. H. L. Hansen and Z. Wang, “Audio anti-spoofing using simple attention module and joint optimization based on additive angular margin loss and meta-learning,” Interspeech , 2022
2022
Later among the works it cites.
J. weon Jung, H.-S. Heo, H. Tak, H. jin Shim, J. S. Chung, B.-J. Lee, H. jin Yu, and N. W. D. Evans, “Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6367–6371, 2022
2022
Later among the works it cites.
R. Yan, C. Wen, S. Zhou, T. Guo, W. Zou, and X. Li, “Audio deepfake detection system with neural stitching for add 2022,” ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 9226–9230, 2022
2022
Later among the works it cites.
H. Wu, H.-C. Kuo, N. Zheng, K.-H. Hung, H. yi Lee, Y. Tsao, H.-M. Wang, and H. M. Meng, “Partially fake audio detection by self-attention-based fake span discovery,” ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 9236–9240, 2022
2022
Later among the works it cites.
C. Wang, J. Yi, J. Tao, H. Sun, X. Chen, Z. Tian, H. Ma, C. Fan, and R. Fu, “Fully automated end-to-end fake audio detection,” Proceedings of the 1st International Workshop on Deepfake Detection for Audio Multimedia , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. H. Kinnunen, M. Todisco, J. Yamagishi, N. W. D. Evans, A. Nautsch, and K.-A. Lee, “Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 2507–2522, 2022
2022
Later among the works it cites.
H. Tak, M. R. Kamble, J. Patino, M. Todisco, and N. W. D. Evans, “Rawboost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,” ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6382–6386, 2022
2022
Later among the works it cites.
N. M. Müller, P. Czempin, F. Dieckmann, A. Froghyar, and K. Böttinger, “Does audio deepfake detection generalize?” in Interspeech , 2022
2022
Later among the works it cites.
J. Yi, J. Tao, R. Fu, X. Yan, C. Wang, T. Wang, C. Y. Zhang, X. Zhang, Y. Zhao, Y. Ren, L. Xu, J. Zhou, H. Gu, Z. Wen, S. Liang, Z. Lian, S. Nie, and H. Li, “Add 2023: the second audio deepfake detection challenge,” 2023
2023
Closest in time.
C. Wang, J. Yi, J. Tao, C. Zhang, S. Zhang, and X. Chen, “Detection of cross-dataset fake audio based on prosodic and pronunciation features,” in Interspeech , 2023
2023
Closest in time.
T.-P. Doan, L. Nguyen-Vu, S. Jung, and K. Hong, “Bts-e: Audio deepfake detection using breathing-talking-silence encoder,” ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023
2023
Closest in time.
J. Kim and S. M. Ban, “Phase-aware spoof speech detection based on res2net with phase network,” ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023
2023
Closest in time.
C. Wang, J. Yi, J. Tao, C. Zhang, S. Zhang, R. Fu, and X. Chen, “To-rawnet: Improving rawnet with tcn and orthogonal regularization for fake audio detection,” in Interspeech , 2023
2023
Closest in time.
X. Liu, M. Liu, L. Wang, K.-A. Lee, H. Zhang, and J. Dang, “Leveraging positional-related local-global dependency for synthetic speech detection,” ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023
2023
Closest in time.
J. Xue, C. Fan, J. Yi, C. Wang, Z. Wen, D. Zhang, and Z. Lv, “Learning from yourself: A self-distillation method for fake speech detection,” ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023
2023
Closest in time.
F. Chen, S. Deng, T. Zheng, Y. He, and J. Han, “Graph-based spectro-temporal dependency modeling for anti-spoofing,” ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023
2023
Closest in time.
S. Ding, Y. Zhang, and Z. Duan, “Samo: Speaker attractor multi-center one-class learning for voice anti-spoofing,” ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023
2023
Closest in time.
X. Zhang, J. Yi, J. Tao, C. Wang, and C. Zhang, “Do you remember? overcoming catastrophic forgetting for fake audio detection,” in Proceedings of the 40-th International Conference on Machine Learning (ICML) , 2023
2023
Closest in time.