Fetching the paper…
Reading the bibliography…
The rapid spread of media content synthesis technology and the potentially damaging impact of audio and video deepfakes on people's lives have raised the need to implement systems able to detect these forgeries automatically.
Pitrelli, J.F., Bakis, R., Eide, E.M., Fernandez, R., Hamza, W., Picheny, M.A.: The IBM expressive text-to-speech synthesis system for American English. IEEE Transactions on Audio, Speech, and Language Processing 14
2006
Earlier work this paper cites.
Busso, C., Bulut, M., Lee, C.C., Kazemzadeh, A., Mower, E., Kim, S., Chang, J.N., Lee, S., Narayanan, S.S.: IEMOCAP: Interactive emotional dyadic motion capture database. Language resources and evaluation 42
2008
Earlier work this paper cites.
Wang, Z.F., Wei, G., He, Q.H.: Channel pattern noise based playback attack detection algorithm for speaker recognition. In: IEEE International Conference on Machine Learning and Cybernetics (ICMLC) (2011)
2011
Earlier work this paper cites.
King, S., Karaiskos, V.: The Blizzard Challenge 2013. In: Blizzard Challenge Workshop (2013)
2013
Earlier work this paper cites.
Panayotov, V., Chen, G., Povey, D., Khudanpur, S.: Librispeech: an ASR corpus based on public domain audio books. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2015)
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016)
2016
Earlier work this paper cites.
Ito, K., Johnson, L.: The LJ Speech Dataset. https://keithito.com/LJ-Speech-Dataset/ (2017)
2017
Earlier work this paper cites.
Nagrani, A., Chung, J.S., Zisserman, A.: VoxCeleb: a large-scale speaker identification dataset. In: Conference of the International Speech Communication Association (INTERSPEECH) (2017)
2017
Earlier work this paper cites.
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R.J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., Le, Q., Agiomyrgiannakis, Y., Clark, R., Saurous, R.A.: Tacotron: Towards end-to-end speech synthesis. In: Conference of the International Speech Communication Association (INTERSPEECH) (2017)
2017
Earlier work this paper cites.
Chung, J.S., Nagrani, A., Zisserman, A.: Voxceleb2: Deep speaker recognition. In: Conference of the International Speech Communication Association (INTERSPEECH) (2018)
2018
Earlier work this paper cites.
Hu, J., Shen, L., Sun, G.: Squeeze-and-excitation networks. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Earlier work this paper cites.
Li, Y., Chang, M.C., Lyu, S.: In ictu oculi: Exposing ai created fake videos by detecting eye blinking. In: IEEE International Workshop on Information Forensics and Security (WIFS) (2018)
2018
Earlier work this paper cites.
Li, Y., Lyu, S.: Exposing deepfake videos by detecting face warping artifacts. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Earlier work this paper cites.
Okabe, K., Koshinaka, T., Shinoda, K.: Attentive statistics pooling for deep speaker embedding. In: Conference of the International Speech Communication Association (INTERSPEECH) (2018)
2018
Earlier work this paper cites.
Skerry-Ryan, R., Battenberg, E., Xiao, Y., Wang, Y., Stanton, D., Shor, J., Weiss, R., Clark, R., Saurous, R.A.: Towards end-to-end prosody transfer for expressive speech synthesis with tacotron. In: International Conference on Machine Learning (ICML) (2018)
2018
Earlier work this paper cites.
Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., Khudanpur, S.: X-vectors: Robust DNN embeddings for speaker recognition. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2018)
2018
Earlier work this paper cites.
Wang, Y., Stanton, D., Zhang, Y., Ryan, R.S., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., Saurous, R.A.: Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis. In: International Conference on Machine Learning (ICML) (2018)
2018
Cited alongside, same era.
Alzantot, M., Wang, Z., Srivastava, M.B.: Deep residual neural networks for audio spoofing detection. In: Conference of the International Speech Communication Association (INTERSPEECH) (2019)
2019
Cited alongside, same era.
Forbes: Deepfakes, revenge porn, and the impact on women. https://www.forbes.com/sites/chenxiwang/2019/11/01/deepfakes-revenge-porn-and-the-impact-on-women/?sh=45b66a961f53
2019
Cited alongside, same era.
Guardian, T.: The rise of the deepfake and the threat to democracy. https://www.theguardian.com/technology/ng-interactive/2019/jun/22/the-rise-of-the-deepfake-and-the-threat-to-democracy
2019
Cited alongside, same era.
Kamble, M.R., Sailor, H.B., Patil, H.A., Li, H.: Advances in anti-spoofing: from the perspective of ASVspoof challenges. APSIPA Transactions on Signal and Information Processing (2020)
2020
Later among the works it cites.
Verdoliva, L.: Media forensics and deepfakes: an overview. IEEE Journal of Selected Topics in Signal Processing 14
2020
Later among the works it cites.
Agarwal, S., Farid, H.: Detecting Deep-Fake Videos From Aural and Oral Dynamics. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Later among the works it cites.
Bonettini, N., Cannas, E.D., Mandelli, S., Bondi, L., Bestagini, P., Tubaro, S.: Video Face Manipulation Detection Through Ensemble of CNNs. In: International Conference on Pattern Recognition (ICPR) (2021)
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lieto, A., Moro, D., Devoti, F., Parera, C., Lipari, V., Bestagini, P., Tubaro, S.: “Hello? Who Am I Talking to?" A Shallow CNN Approach for Human vs. Bot Speech Classification. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2019)
2019
Cited alongside, same era.
Malik, H.: Securing voice-driven interfaces against fake (cloned) audio attacks. In: IEEE Conference on Multimedia Information Processing and Retrieval (MIPR) (2019)
2019
Cited alongside, same era.
Snyder, D., Garcia-Romero, D., Sell, G., McCree, A., Povey, D., Khudanpur, S.: Speaker recognition for multi-speaker conversations using x-vectors. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2019)
2019
Cited alongside, same era.
Todisco, M., Wang, X., Vestman, V., Sahidullah, M., Delgado, H., Nautsch, A., Yamagishi, J., Evans, N., Kinnunen, T., Lee, K.A.: ASVspoof 2019: Future horizons in spoofed and fake audio detection. In: Conference of the International Speech Communication Association (INTERSPEECH) (2019)
2019
Cited alongside, same era.
Westerlund, M.: The emergence of deepfake technology: A review. Technology Innovation Management Review 9
2019
Cited alongside, same era.
Yang, X., Li, Y., Lyu, S.: Exposing deep fakes using inconsistent head poses. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2019)
2019
Cited alongside, same era.
Zeinali, H., Wang, S., Silnova, A., Matějka, P., Plchot, O.: BUT system description to voxceleb speaker recognition challenge 2019. In: The VoxCeleb Challenge Workshop (2019)
2019
Cited alongside, same era.
Zhang, X., Karaman, S., Chang, S.F.: Detecting and simulating artifacts in gan fake images. In: IEEE International Workshop on Information Forensics and Security (WIFS) (2019)
2019
Cited alongside, same era.
2021
Later among the works it cites.
Cozzolino, D., Rössler, A., Thies, J., Nießner, M., Verdoliva, L.: Id-reveal: Identity-aware deepfake video detection. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Later among the works it cites.
Forbes: Fraudsters Cloned Company Director’s Voice In 35$ Million Bank Heist, Police Find. https://www.forbes.com/sites/thomasbrewster/2021/10/14/huge-bank-fraud-uses-deep-fake-voice-tech-to-steal-millions
2021
Later among the works it cites.
Gao, Y., Vuong, T., Elyasi, M., Bharaj, G., Singh, R.: Generalized Spoofing Detection Inspired from Audio Generation Artifacts. In: Conference of the International Speech Communication Association (INTERSPEECH) (2021)
2021
Later among the works it cites.
Hosler, B., Salvi, D., Murray, A., Antonacci, F., Bestagini, P., Tubaro, S., Stamm, M.C.: Do Deepfakes Feel Emotions? A Semantic Approach to Detecting Deepfakes via Emotional Inconsistencies. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
de Ruiter, A.: The distinct wrong of deepfakes. Philosophy & Technology 34
2021
Later among the works it cites.
Tak, H., Patino, J., Todisco, M., Nautsch, A., Evans, N., Larcher, A.: End-to-end anti-spoofing with RawNet2. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2021)
2021
Later among the works it cites.
The New York Times: Pennsylvania Woman Accused of Using Deepfake Technology to Harass Cheerleaders. https://www.nytimes.com/2021/03/14/us/raffaela-spone-victory-vipers-deepfake.html
2021
Later among the works it cites.
Yamagishi, J., Wang, X., Todisco, M., Sahidullah, M., Patino, J., Nautsch, A., Liu, X., Lee, K.A., Kinnunen, T., Evans, N., et al.: ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection. In: Automatic Speaker Verification and Spoofing Countermeasures Challenge (2021)
2021
Later among the works it cites.
Conti, E., Salvi, D., Borrelli, C., Hosler, B., Bestagini, P., Antonacci, F., Sarti, A., Stamm, M.C., Tubaro, S.: Deepfake Speech Detection Through Emotion Recognition: a Semantic Approach. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2022)
2022
Closest in time.