Fetching the paper…
Reading the bibliography…
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods.
Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems , Vol. 33. Curran Associates, Inc., 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1982–1991
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu. 2023b · 1991
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. 2004 · 2004
Earlier work this paper cites.
The DeepFake Detection Challenge (DFDC) Dataset
Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer. 2020 · 2006
Earlier work this paper cites.
Converting video formats with FFmpeg
Suramya Tomar. 2006 · 2006
Earlier work this paper cites.
Attentron: Few-Shot Text-to-Speech Utilizing Attention-Based Variable-Length Embedding. In Interspeech 2020 . ISCA, 2007–2011
Seungwoo Choi, Seungju Han, Dongyoung Kim, and Sungjoo Ha. 2020 · 2011
Earlier work this paper cites.
Xception: Deep Learning With Depthwise Separable Convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1251–1258
Francois Chollet. 2017 · 2017
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Advances in Neural Information Processing Systems , Vol. 30. Curran Associates, Inc
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
Temporal Action Detection With Structured Segment Networks. In Proceedings of the IEEE International Conference on Computer Vision . 2914–2923
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin. 2017 · 2017
Earlier work this paper cites.
MesoNet: a Compact Facial Video Forgery Detection Network. In 2018 IEEE International Workshop on Information Forensics and Security (WIFS) . 1–7
Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. 2018 · 2018
Earlier work this paper cites.
VoxCeleb2: Deep Speaker Recognition. In Interspeech 2018 . ISCA, 1086–1090
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (NIPS’18) . Curran Associates Inc., Red Hook, NY, USA, 4485–4495
Ye Jia, Yu Zhang, Ron J. Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, and Yonghui Wu. 2018 · 2018
Earlier work this paper cites.
DeepFakes: a New Threat to Face Recognition? Assessment and Detection
Pavel Korshunov and Sebastien Marcel. 2018 · 2018
Earlier work this paper cites.
Generalized End-to-End Loss for Speaker Verification. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 4879–4883
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno. 2018 · 2018
Earlier work this paper cites.
Hierarchical Cross-Modal Talking Face Generation With Dynamic Pixel-Wise Loss. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7832–7841
Lele Chen, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu. 2019 · 2019
Earlier work this paper cites.
Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. 2019 · 2019
Earlier work this paper cites.
Exposing DeepFake Videos By Detecting Face Warping Artifacts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops . 7
Yuezun Li and Siwei Lyu. 2019 · 2019
Earlier work this paper cites.
Contributing Data to Deepfake Detection Research
Dufou Nick and Jigsaw Andrew. 2019 · 2019
Earlier work this paper cites.
FaceForensics++: Learning to Detect Manipulated Facial Images. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 1–11
Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Niessner. 2019 · 2019
Earlier work this paper cites.
Exposing Deep Fakes Using Inconsistent Head Poses. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 8261–8265
Xin Yang, Yuezun Li, and Siwei Lyu. 2019 · 2019
Earlier work this paper cites.
Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and Localization. In Proceedings of the 28th ACM International Conference on Multimedia (MM ’20) . Association for Computing Machinery, New York, NY, USA, 439–447
Komal Chugh, Parul Gupta, Abhinav Dhall, and Ramanathan Subramanian. 2020 · 2020
Earlier work this paper cites.
Real Time Speech Enhancement in the Waveform Domain. In Interspeech 2020 . Shanghai, China, 3291–3295
Alexandre Défossez, Gabriel Synnaeve, and Yossi Adi. 2020 · 2020
Earlier work this paper cites.
DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2889–2898
Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. 2020 · 2020
Cited alongside, same era.
Emotions Don’t Lie: An Audio-Visual Deepfake Detection Method using Affective Cues. In Proceedings of the 28th ACM International Conference on Multimedia (MM ’20) . Association for Computing Machinery, New York, NY, USA, 2823–2832
Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha. 2020 · 2020
Cited alongside, same era.
A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild. In Proceedings of the 28th ACM International Conference on Multimedia (MM ’20) . Association for Computing Machinery, New York, NY, USA, 484–492
K R Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, and C.V. Jawahar. 2020 · 2020
Cited alongside, same era.
Thinking in Frequency: Face Forgery Detection by Mining Frequency-Aware Clues. In Proceedings of the European Conference on Computer Vision (ECCV) (Lecture Notes in Computer Science) , Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.). Springer International Publishing, Cham, 86–103
Make-A-Video: Text-to-Video Generation without Text-Video Data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. 2022 · 2022
Later among the works it cites.
M2TR: Multi-modal Multi-scale Transformers for Deepfake Detection. In Proceedings of the 2022 International Conference on Multimedia Retrieval (ICMR ’22) . Association for Computing Machinery, New York, NY, USA, 615–623
Junke Wang, Zuxuan Wu, Wenhao Ouyang, Xintong Han, Jingjing Chen, Yu-Gang Jiang, and Ser-Nam Li. 2022b · 2022
Later among the works it cites.
ADD 2022: the First Audio Deep Synthesis Detection Challenge
Jiangyan Yi, Ruibo Fu, Jianhua Tao, Shuai Nie, Haoxin Ma, Chenglong Wang, Tao Wang, Zhengkun Tian, Ye Bai, Cunhang Fan, Shan Liang, Shiming Wang, Shuai Zhang, Xinrui Yan, Le Xu, Zhengqi Wen, Haizhou Li, Zheng Lian, and Bin Liu. 2022 · 2022
Later among the works it cites.
ActionFormer: Localizing Moments of Actions with Transformers. In Proceedings of the European Conference on Computer Vision (ECCV) (Lecture Notes in Computer Science) , Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (Eds.). Springer Nature Switzerland, Cham, 492–510
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. 2020 · 2020
Cited alongside, same era.
WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection. In Proceedings of the 28th ACM International Conference on Multimedia (MM ’20) . Association for Computing Machinery, New York, NY, USA, 2382–2390
Bojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma, and Yu-Gang Jiang. 2020 · 2020
Cited alongside, same era.
SC-GlowTTS: An Efficient Zero-Shot Multi-Speaker Text-To-Speech Model. In Interspeech 2021 . ISCA, 3645–3649
Edresson Casanova, Christopher Shulby, Eren Gölge, Nicolas Michael Müller, Frederico Santos De Oliveira, Arnaldo Candido Jr., Anderson Da Silva Soares, Sandra Maria Aluisio, and Moacir Antonelli Ponti. 2021 · 2021
Cited alongside, same era.
AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 5784–5794
Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang. 2021 · 2021
Cited alongside, same era.
Lips Don’t Lie: A Generalisable and Robust Approach To Face Forgery Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5039–5049
Alexandros Haliassos, Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. 2021 · 2021
Cited alongside, same era.
ForgeryNet: A Versatile Benchmark for Comprehensive Forgery Analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4360–4369
Yinan He, Bei Gan, Siyu Chen, Yichun Zhou, Guojun Yin, Luchuan Song, Lu Sheng, Jing Shao, and Ziwei Liu. 2021 · 2021
Cited alongside, same era.
FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset
Hasam Khalid, Shahroz Tariq, and Simon S. Woo. 2021 · 2021
Cited alongside, same era.
Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech. In Proceedings of the 38th International Conference on Machine Learning . PMLR, 5530–5540
Jaehyeon Kim, Jungil Kong, and Juhee Son. 2021 · 2021
Cited alongside, same era.
KoDF: A Large-Scale Korean DeepFake Detection Dataset. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 10744–10753
Patrick Kwon, Jaeseong You, Gyuhyeon Nam, Sungwoo Park, and Gyeongsu Chae. 2021 · 2021
Cited alongside, same era.
Chen-Lin Zhang, Jianxin Wu, and Yin Li. 2022 · 2022
Later among the works it cites.
Audio-Visual Face Reenactment. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 5178–5187
Madhav Agarwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, and C. V. Jawahar. 2023 · 2023
Closest in time.
Glitch in the matrix: A large scale benchmark for content driven audio–visual forgery detection and localization
Zhixi Cai, Shreya Ghosh, Abhinav Dhall, Tom Gedeon, Kalin Stefanov, and Munawar Hayat. 2023a · 2023
Closest in time.
Self-Supervised Video Forensics by Audio-Visual Anomaly Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10491–10503
Chao Feng, Ziyang Chen, and Andrew Owens. 2023 · 2023
Closest in time.
DeepFake detection algorithm based on improved vision transformer
Young-Jin Heo, Woon-Ha Yeo, and Byung-Gyu Kim. 2023 · 2023
Closest in time.
AVFakeNet: A unified end-to-end Dense Swin Transformer deep learning model for audio–visual deepfakes detection
Hafsa Ilyas, Ali Javed, and Khalid Mahmood Malik. 2023 · 2023
Closest in time.
ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild
Xuechen Liu, Xin Wang, Md Sahidullah, Jose Patino, Héctor Delgado, Tomi Kinnunen, Massimiliano Todisco, Junichi Yamagishi, Nicholas Evans, Andreas Nautsch, and Kong Aik Lee. 2023 · 2023
Closest in time.
DF-Platter: Multi-Face Heterogeneous Deepfake Dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9739–9748
Kartik Narayan, Harsh Agarwal, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, and Richa Singh. 2023 · 2023
Closest in time.
Powerset multi-class cross entropy loss for neural speaker diarization. In INTERSPEECH 2023 . ISCA, 3222–3226
Alexis Plaquet and Hervé Bredin. 2023 · 2023
Closest in time.
Robust Speech Recognition via Large-Scale Weak Supervision. In Proceedings of the 40th International Conference on Machine Learning . PMLR, 28492–28518
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine Mcleavey, and Ilya Sutskever. 2023 · 2023
Closest in time.
Multimodaltrace: Deepfake Detection Using Audiovisual Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 993–1000
Muhammad Anas Raza and Khalid Mahmood Malik. 2023 · 2023
Closest in time.
TriDet: Temporal Action Detection With Relative Boundary Modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18857–18866
Dingfeng Shi, Yujie Zhong, Qiong Cao, Lin Ma, Jia Li, and Dacheng Tao. 2023 · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
PVASS-MDD: Predictive Visual-audio Alignment Self-supervision for Multimodal Deepfake Detection
Yang Yu, Xiaolong Liu, Rongrong Ni, Siyuan Yang, Yao Zhao, and Alex C. Kot. 2023 · 2023
Closest in time.
AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 27102–27112
Trevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki, Ben Colman, Yaser Yacoob, Ali Shahriyari, and Gaurav Bharaj. 2024 · 2024
Closest in time.
Detecting and Grounding Multi-Modal Media Manipulation and Beyond
Rui Shao, Tianxing Wu, Jianlong Wu, Liqiang Nie, and Ziwei Liu. 2024 · 2024
Closest in time.
Exploiting Modality-Specific Features for Multi-Modal Manipulation Detection and Grounding. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 4935–4939
Jiazhen Wang, Bin Liu, Changtao Miao, Zhiwei Zhao, Wanyi Zhuang, Qi Chu, and Nenghai Yu. 2024 · 2024
Closest in time.
AVoiD-DF: Audio-Visual Joint Learning for Detecting Deepfake
Wenyuan Yang, Xiaoyu Zhou, Zhikai Chen, Bofei Guo, Zhongjie Ba, Zhihua Xia, Xiaochun Cao, and Kui Ren. 2023 · 2029
Closest in time.