Fetching the paper…
Reading the bibliography…
Thanks to recent advances in deep learning, sophisticated generation tools exist, nowadays, that produce extremely realistic synthetic speech.
A. Janicki, “Spoofing countermeasure based on analysis of linear prediction error,” in Sixteenth Annual Conference of the International Speech Communication Association , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for environmental sound classification,” in ACM international conference on Multimedia , 2015, pp. 1015–1018
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 5220–5224
2017
Earlier work this paper cites.
W. Cai, J. Chen, and M. Li, “Exploring the Encoding Layer and Loss Function in End-to-End Speaker and Language Recognition System,” in The Speaker and Language Recognition Workshop (Odyssey) , 2018, pp. 74–81
2018
Earlier work this paper cites.
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive Statistics Pooling for Deep Speaker Embedding,” in Interspeech , 2018, pp. 2252–2256
2018
Earlier work this paper cites.
S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, “Cbam: Convolutional block attention module,” in European Conference on Computer Vision (ECCV) , 2018, pp. 3–19
2018
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep Speaker Recognition,” in Interspeech , 2018, pp. 1086–1090
2018
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
T. Kinnunen, K. A. Lee, H. Delgado, N. Evans, M. Todisco, M. Sahidullah, J. Yamagishi, and D. A. Reynolds, “t-DCF: a detection cost function for the tandem assessment of spoofing countermeasures and automatic speaker verification,” in Odyssey , 2018, pp. 312–319
2018
Earlier work this paper cites.
E. AlBadawy, S. Lyu, and H. Farid, “Detecting AI-Synthesized Speech using Bispectral Analysis,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , 2019
2019
Earlier work this paper cites.
S. Agarwal, H. Farid, Y. Gu, M. He, K. Nagano, and H. Li, “Protecting world leaders against deep fakes,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , June 2019
2019
Cited alongside, same era.
J. Wang, K.-C. Wang, M. T. Law, F. Rudzicz, and M. Brudno, “Centroid-based deep metric learning for speaker recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 3652–3656
2019
Cited alongside, same era.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “ArcFace: Additive Angular Margin Loss for Deep Face Recognition,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 4690–4699
2019
Cited alongside, same era.
G. Lavrentyeva, S. Novoselov, A. Tseren, M. Volkova, A. Gorlanov, and A. Kozlov, “STC Antispoofing Systems for the ASVspoof2019 Challenge,” in Interspeech , 2019, pp. 1033–1037
2019
Cited alongside, same era.
X. Wang and J. Yamagishi, “A Comparative Study on Recent Neural Spoofing Countermeasures for Synthetic Speech Detection,” in Interspeech , 2021, pp. 4259–4263
2021
Later among the works it cites.
Z. Zhang, X. Yi, and X. Zhao, “Fake speech detection using residual network with transformer encoder,” in ACM Workshop on Information Hiding and Multimedia Security (IH&MMSec) , 2021, p. 13–22
2021
Later among the works it cites.
H. Tak, J.-w. Jung, J. Patino, M. Kamble, M. Todisco, and N. Evans, “End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection,” in Automatic Speaker Verification and Spoofing Countermeasures Challenge , 2021
2021
Later among the works it cites.
C. Borrelli, P. Bestagini, F. Antonacci, A. Sarti, and S. Tubaro, “Synthetic speech detection through short-term and long-term prediction traces,” EURASIP Journal on Information Security , vol. 2021, no. 1, pp. 1–14, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Yamagishi, M. Todisco, M. Sahidullah, H. Delgado, X. Wang, N. Evans, T. Kinnunen, K. A. Lee, V. Vestman, and A. Nautsch, “Asvspoof 2019: the automatic speaker verification spoofing and countermeasures challenge evaluation plan.” 2019. [Online]. Available: http://www.asvspoof.org/asvspoof2019/asvspoof2019_evaluation_plan.pdf
2019
Cited alongside, same era.
R. K. Das, X. Tian, T. Kinnunen, and H. Li, “The attacker’s perspective on automatic speaker verification: An overview,” in Interspeech , 2020, pp. 4213–4217
2020
Cited alongside, same era.
S. Agarwal, H. Farid, T. El-Gaaly, and S.-N. Lim, “Detecting deep-fake videos from appearance and behavior,” in IEEE International Workshop on Information Forensics and Security (WIFS) , 2020, pp. 1–6
2020
Cited alongside, same era.
J. S. Chung, J. Huh, S. Mun, M. Lee, H.-S. Heo, S. Choe, C. Ham, S. Jung, B.-J. Lee, and I. Han, “In Defence of Metric Learning for Speaker Recognition,” in Interspeech , 2020, pp. 2977–2981
2020
Cited alongside, same era.
2020
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,” in Interspeech , 2020, pp. 3830–3834
2020
Cited alongside, same era.
P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan, “Supervised contrastive learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 18 661–18 673, 2020
2020
Cited alongside, same era.
X. Wang, J. Yamagishi, M. Todisco, H. Delgado, A. Nautsch, N. Evans, M. Sahidullah, V. Vestman, T. Kinnunen, K. A. Lee et al. , “ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech,” Computer Speech & Language , vol. 64, p. 101114, 2020
2020
Cited alongside, same era.
H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, “End-to-End anti-spoofing with RawNet2,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 6369–6373
2021
Later among the works it cites.
Y. Zhang, F. Jiang, and Z. Duan, “One-class learning towards synthetic voice spoofing detection,” IEEE Signal Processing Letters , vol. 28, pp. 937–941, 2021
2021
Later among the works it cites.
D. Cozzolino, A. Rössler, J. Thies, M. Nießner, and L. Verdoliva, “ID-Reveal: Identity-aware DeepFake Video Detection,” in IEEE International Conference on Computer Vision (ICCV) , October 2021
2021
Later among the works it cites.
H. Khalid, S. Tariq, M. Kim, and S. S. Woo, “FakeAVCeleb: A novel audio-video multimodal deepfake dataset,” in Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) , 2021
2021
Later among the works it cites.
2022
Closest in time.
D. Castan, M. H. Rahman, S. Bakst, C. Cobo-Kroenke, M. McLaren, M. Graciarena, and A. Lawson, “Speaker-targeted synthetic speech detection,” in The Speaker and Language Recognition Workshop , 2022
2022
Closest in time.
N. M. Müller, P. Czempin, F. Dieckmann, A. Froghyar, and K. Böttinger, “Does Audio Deepfake Detection Generalize?” in Interspeech , 2022, pp. 2783–2787
2022
Closest in time.
M. Sahidullah, T. Kinnunen, and C. Hanilçi, “A comparison of features for synthetic speech detection,” in Interspeech , 2015, pp. 2087–2091
2091
Closest in time.