Fetching the paper…
Reading the bibliography…
Deepfake technology has rapidly advanced and poses significant threats to information integrity and trust in online multimedia.
Rational decisions
Irving John Good · 1952
Earlier work this paper cites.
Visual contribution to speech intelligibility in noise
William H Sumby and Irwin Pollack · 1954
Earlier work this paper cites.
Visual illusion induced by sound
Ladan Shams, Yukiyasu Kamitani, and Shinsuke Shimojo · 2002
Earlier work this paper cites.
Seeing what you hear: Cross-modal illusions and perception
Casey O’Callaghan · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Transitions in neural oscillations reflect prediction errors generated in audiovisual speech
Luc H Arnal, Valentin Wyart, and Anne-Lise Giraud · 2011
Earlier work this paper cites.
When what you see is not what you hear
C. Chandrasekaran and A. Ghazanfar · 2011
Earlier work this paper cites.
A neural basis for interindividual differences in the mcgurk effect, a multisensory speech illusion
Audrey R Nath and Michael S Beauchamp · 2012
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
The contribution of visual information to the perception of speech in noise with and without informative temporal fine structure
Paula C Stacey, Pádraig T Kitterick, Saffron D Morris, and Christian J Sumner · 2016
Earlier work this paper cites.
Phoneme-to-viseme mappings: the good, the bad, and the ugly
Helen L Bear and Richard Harvey · 2017
Earlier work this paper cites.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Mesonet: a compact facial video forgery detection network
Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen · 2018
Earlier work this paper cites.
Deep audio-visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2018
Earlier work this paper cites.
Lrs3-ted: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
What you see depends on what you hear: Temporal averaging and crossmodal integration
Lihan Chen, Xiaolin Zhou, Hermann J Müller, and Zhuanghua Shi · 2018
Earlier work this paper cites.
VoxCeleb2: Deep speaker recognition
J Chung, A Nagrani, and A Zisserman · 2018
Earlier work this paper cites.
Deepfakes: a new threat to face recognition? assessment and detection
Pavel Korshunov and Sébastien Marcel · 2018
Earlier work this paper cites.
Detection of fake images via the ensemble of deep representations from multi color spaces
Peisong He, Haoliang Li, and Hongxia Wang · 2019
Earlier work this paper cites.
Bmn: Boundary-matching network for temporal action proposal generation
Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen · 2019
Earlier work this paper cites.
Faceforensics++: Learning to detect manipulated facial images
Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner · 2019
Earlier work this paper cites.
Exposing deep fakes using inconsistent head poses
Xin Yang, Yuezun Li, and Siwei Lyu · 2019
Earlier work this paper cites.
Not made for each other-audio-visual dissonance-based deepfake detection and localization
Komal Chugh, Parul Gupta, Abhinav Dhall, and Ramanathan Subramanian · 2020
Earlier work this paper cites.
The deepfake detection challenge (dfdc) dataset
Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer · 2020
Earlier work this paper cites.
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al · 2020
Earlier work this paper cites.
Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection
Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy · 2020
Cited alongside, same era.
Face x-ray for more general face forgery detection
Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo · 2020
Cited alongside, same era.
Celeb-df: A large-scale challenging dataset for deepfake forensics
Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu · 2020
Cited alongside, same era.
Emotions don’t lie: An audio-visual deepfake detection method using affective cues
Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha · 2020
Cited alongside, same era.
Thinking in frequency: Face forgery detection by mining frequency-aware clues
Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao · 2020
Cited alongside, same era.
Distance-iou loss: Faster and better learning for bounding box regression
Audio-visual person-of-interest deepfake detection
Davide Cozzolino, Alessandro Pianese, Matthias Nießner, and Luisa Verdoliva · 2023
Later among the works it cites.
Self-supervised video forensics by audio-visual anomaly detection
Chao Feng, Ziyang Chen, and Andrew Owens · 2023
Later among the works it cites.
Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio-visual deepfakes detection
Hafsa Ilyas, Ali Javed, and Khalid Mahmood Malik · 2023
Later among the works it cites.
Spatio-temporal catcher: A self-supervised transformer for deepfake video detection
Maosen Li, Xurong Li, Kun Yu, Cheng Deng, Heng Huang, Feng Mao, Hui Xue, and Minghao Li · 2023
Later among the works it cites.
Deepfakes generation and detection: State-of-the-art, open challenges, countermeasures, and way forward
Momina Masood, Mariam Nawaz, Khalid Mahmood Malik, Ali Javed, Aun Irtaza, and Hafiz Malik · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhaohui Zheng, Ping Wang, Wei Liu, Jinze Li, Rongguang Ye, and Dongwei Ren · 2020
Cited alongside, same era.
Wilddeepfake: A challenging real-world dataset for deepfake detection
Bojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma, and Yu-Gang Jiang · 2020
Cited alongside, same era.
Hear me out: Fusional approaches for audio augmented temporal action localization
Anurag Bagchi, Jazib Mahmood, Dolton Fernandes, and Ravi Kiran Sarvadevabhatla · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Cited alongside, same era.
Deepfake video detection using audio-visual consistency
Yewei Gu, Xianfeng Zhao, Chen Gong, and Xiaowei Yi · 2021
Cited alongside, same era.
Lips don’t lie: A generalisable and robust approach to face forgery detection
Alexandros Haliassos, Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2021
Cited alongside, same era.
Fakeavceleb: A novel audio-video multimodal deepfake dataset
Hasam Khalid, Shahroz Tariq, Minha Kim, and Simon S Woo · 2021
Cited alongside, same era.
Df-platter: Multi-face heterogeneous deepfake dataset
Kartik Narayan, Harsh Agarwal, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, and Richa Singh · 2023
Later among the works it cites.
Powerset multi-class cross entropy loss for neural speaker diarization
Alexis Plaquet and Hervé Bredin · 2023
Later among the works it cites.
Multimodaltrace: Deepfake detection using audiovisual representation learning
Muhammad Anas Raza and Khalid Mahmood Malik · 2023
Later among the works it cites.
Detecting deepfakes without seeing any
Tal Reiss, Bar Cavia, and Yedid Hoshen · 2023
Later among the works it cites.
Tridet: Temporal action detection with relative boundary modeling
Dingfeng Shi, Yujie Zhong, Qiong Cao, Lin Ma, Jia Li, and Dacheng Tao · 2023
Later among the works it cites.
Locate and verify: A two-stream network for improved deepfake detection
Chao Shuai, Jieming Zhong, Shuang Wu, Feng Lin, Zhibo Wang, Zhongjie Ba, Zhenguang Liu, Lorenzo Cavallaro, and Kui Ren · 2023
Later among the works it cites.
Learning on gradients: Generalized artifacts representation for gan-generated images detection
Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, and Yunchao Wei · 2023
Later among the works it cites.
Videomae v2: Scaling video masked autoencoders with dual masking
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao · 2023
Later among the works it cites.
Avoid-df: Audio-visual joint learning for detecting deepfake
Wenyuan Yang, Xiaoyu Zhou, Zhikai Chen, Bofei Guo, Zhongjie Ba, Zhihua Xia, Xiaochun Cao, and Kui Ren · 2023
Later among the works it cites.
Pvass-mdd: predictive visual-audio alignment self-supervision for multimodal deepfake detection
Yang Yu, Xiaolong Liu, Rongrong Ni, Siyuan Yang, Yao Zhao, and Alex C Kot · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Hang Zhang, Xin Li, and Lidong Bing · 2023
Later among the works it cites.
Ummaformer: A universal multimodal-adaptive transformer framework for temporal forgery localization
Rui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu, Yang Zhou, and Qiang Zeng · 2023
Later among the works it cites.
Lost in translation: Lip-sync deepfake detection from audio-video mismatch
Matyas Bohacek and Hany Farid · 2024
Closest in time.
Av-deepfake1m: A large-scale llm-driven audio-visual deepfake dataset
Zhixi Cai, Shreya Ghosh, Aman Pankaj Adatia, Munawar Hayat, Abhinav Dhall, Tom Gedeon, and Kalin Stefanov · 2024
Closest in time.
Contextual cross-modal attention for audio-visual deepfake detection and localization
Vinaya Sree Katamneni and Ajita Rattani · 2024
Closest in time.
Zero-shot fake video detection by audio-visual consistency
Xiaolou Li, Zehua Liu, Chen Chen, Lantian Li, Li Guo, and Dong Wang · 2024
Closest in time.
Frade: Forgery-aware audio-distilled multimodal learning for deepfake detection
Fan Nie, Jiangqun Ni, Jian Zhang, Bin Zhang, and Weizhe Zhang · 2024
Closest in time.
Avff: Audio-visual feature fusion for video deepfake detection
Trevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki, Ben Colman, Yaser Yacoob, Ali Shahriyari, and Gaurav Bharaj · 2024
Closest in time.
Deepfake generation and detection: A benchmark and survey
Gan Pei, Jiangning Zhang, Menghan Hu, Guangtao Zhai, Chengjie Wang, Zhenyu Zhang, Jian Yang, Chunhua Shen, and Dacheng Tao · 2024
Closest in time.
Building robust video-level deepfake detection via audio-visual local-global interactions
Yifan Wang, Xuecheng Wu, Jia Zhang, Mohan Jing, Keda Lu, Jun Yu, Wen Su, Fang Gao, Qingsong Liu, Jianqing Sun, et al · 2024
Closest in time.
Audio-visual deepfake detection using articulatory representation learning
Yujia Wang and Hua Huang · 2024
Closest in time.
Joint audio-visual attention with contrastive learning for more general deepfake detection
Yibo Zhang, Weiguo Lin, and Junfeng Xu · 2024
Closest in time.
Cross-modality and within-modality regularization for audio-visual deepfake detection
Heqing Zou, Meng Shen, Yuchen Hu, Chen Chen, Eng Siong Chng, and Deepu Rajan · 2024
Closest in time.