Fetching the paper…
Reading the bibliography…
The escalating quality of video generated by advanced video generation methods results in new security challenges, while there have been few relevant research efforts: 1) There is no open-source dataset for generated video detection, 2) No generated video detection method has been proposed so far.
“Visualizing data using t-sne.,”
Laurens Van der Maaten and Geoffrey Hinton, · 2008
Earlier work this paper cites.
“Collecting highly parallel data for paraphrase evaluation,”
David Chen and William B Dolan, · 2011
Earlier work this paper cites.
“Msr-vtt: A large video description dataset for bridging video and language,”
Jun Xu, Tao Mei, Ting Yao, and Yong Rui, · 2016
Earlier work this paper cites.
“Quo vadis, action recognition? a new model and the kinetics dataset,”
Joao Carreira and Andrew Zisserman, · 2017
Earlier work this paper cites.
“Grad-cam: Visual explanations from deep networks via gradient-based localization,”
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra, · 2017
Earlier work this paper cites.
“Slowfast networks for video recognition,”
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He, · 2019
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Earlier work this paper cites.
“Denoising diffusion implicit models,”
Jiaming Song, Chenlin Meng, and Stefano Ermon, · 2020
Earlier work this paper cites.
“Cnn-generated images are surprisingly easy to spot… for now,”
Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A Efros, · 2020
Earlier work this paper cites.
“Leveraging frequency analysis for deep fake image recognition,”
Joel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer, Dorothea Kolossa, and Thorsten Holz, · 2020
Earlier work this paper cites.
“X3d: Expanding architectures for efficient video recognition,”
Christoph Feichtenhofer, · 2020
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al., · 2020
Earlier work this paper cites.
“Improved denoising diffusion probabilistic models,”
Alexander Quinn Nichol and Prafulla Dhariwal, · 2021
Earlier work this paper cites.
“Multiscale vision transformers,”
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer, · 2021
Earlier work this paper cites.
“Taming transformers for high-resolution image synthesis,”
Patrick Esser, Robin Rombach, and Bjorn Ommer, · 2021
Earlier work this paper cites.
“High-resolution image synthesis with latent diffusion models,”
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer, · 2022
Cited alongside, same era.
“Hierarchical text-conditional image generation with clip latents,”
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen, · 2022
Cited alongside, same era.
“Make-a-video: Text-to-video generation without text-video data,”
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al., · 2022
Cited alongside, same era.
“Detecting deepfakes with self-blended images,”
Kaede Shiohara and Toshihiko Yamasaki, · 2022
Cited alongside, same era.
“Frozen clip models are efficient video learners,”
Ziyi Lin, Shijie Geng, Renrui Zhang, Peng Gao, Gerard De Melo, Xiaogang Wang, Jifeng Dai, Yu Qiao, and Hongsheng Li, · 2022
Cited alongside, same era.
“Towards universal fake image detectors that generalize across generative models,”
Utkarsh Ojha, Yuheng Li, and Yong Jae Lee, · 2023
Later among the works it cites.
“Ucf: Uncovering common features for generalizable deepfake detection,”
Zhiyuan Yan, Yong Zhang, Yanbo Fan, and Baoyuan Wu, · 2023
Later among the works it cites.
“Deepfakebench: A comprehensive benchmark of deepfake detection,”
Zhiyuan Yan, Yong Zhang, Xinhang Yuan, Siwei Lyu, and Baoyuan Wu, · 2023
Later among the works it cites.
“Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation,”
Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou, · 2023
Later among the works it cites.
“Deep image fingerprint: Accurate and low budget synthetic image detector,”
Sergey Sinitsa and Ohad Fried, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Zero-shot image-to-image translation,”
Gaurav Parmar, Krishna Kumar Singh, Richard Zhang, Yijun Li, Jingwan Lu, and Jun-Yan Zhu, · 2023
Cited alongside, same era.
“Align your latents: High-resolution video synthesis with latent diffusion models,”
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis, · 2023
Cited alongside, same era.
“Show-1: Marrying pixel and latent diffusion models for text-to-video generation,”
David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao, Lingmin Ran, Yuchao Gu, Difei Gao, and Mike Zheng Shou, · 2023
Cited alongside, same era.
“A survey of detection and mitigation for fake images on social media platforms,”
Dilip Kumar Sharma, Bhuvanesh Singh, Saurabh Agarwal, Lalit Garg, Cheonshik Kim, and Ki-Hyun Jung, · 2023
Cited alongside, same era.
“Identifying and mitigating the security risks of generative ai,”
Clark Barrett, Brad Boyd, Elie Bursztein, Nicholas Carlini, Brad Chen, Jihye Choi, Amrita Roy Chowdhury, Mihai Christodorescu, Anupam Datta, Soheil Feizi, et al., · 2023
Cited alongside, same era.
“Text2video-zero: Text-to-image diffusion models are zero-shot video generators,”
Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi, · 2023
Cited alongside, same era.
“Modelscope text-to-video technical report,”
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang, · 2023
Cited alongside, same era.
“Learning on gradients: Generalized artifacts representation for gan-generated images detection,”
Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, and Yunchao Wei, · 2023
Later among the works it cites.
“Pika,” https://www.pika.art/
Pika Labs, · 2023
Later among the works it cites.
“Gen-2,” https://research.runwayml.com/gen2
Runway Research, · 2023
Later among the works it cites.
“Sora,” https://openai.com/index/sora
OpenAI, · 2024
Closest in time.
“Veo,” https://deepmind.google/technologies/veo/
Google Deepmind, · 2024
Closest in time.
“Kling,” https://kling.kuaishou.com/
Kwai, · 2024
Closest in time.
“zero,” https://huggingface.co/cerspense/zeroscope_v2_576w
Alibaba DAMO Academy for Discovery, · 2024
Closest in time.
“Df40: Toward next-generation deepfake detection,”
Zhiyuan Yan, Taiping Yao, Shen Chen, Yandan Zhao, Xinghe Fu, Junwei Zhu, Donghao Luo, Li Yuan, Chengjie Wang, Shouhong Ding, et al., · 2024
Closest in time.
“Transcending forgery specificity with latent space augmentation for generalizable deepfake detection,”
Zhiyuan Yan, Yuhao Luo, Siwei Lyu, Qingshan Liu, and Baoyuan Wu, · 2024
Closest in time.