Fetching the paper…
Reading the bibliography…
AI-generated video generation continues its journey through the uncanny valley to produce content that is increasingly perceptually indistinguishable from reality.
Support-vector networks
Corinna Cortes and Vladimir Naumovich Vapnik · 1995
Earlier work this paper cites.
Seeing biological motion
Peter Neri, M Concetta Morrone, and David C Burr · 1998
Earlier work this paper cites.
Attacks on digital watermarks: Classification, estimation based attacks, and benchmarks
Sviatoslav Voloshynovskiy, Shelby Pereira, Thierry Pun, Joachim J Eggers, and Jonathan K Su · 2001
Earlier work this paper cites.
Render me real? Investigating the effect of render style on the perception of animated virtual humans
Rachel McDonnell, Martin Breidt, and Heinrich H Bülthoff · 2012
Earlier work this paper cites.
Synthesizing Obama: Learning lip sync from audio
Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman · 2017
Earlier work this paper cites.
FSGAN: Subject agnostic face swapping and reenactment
Yuval Nirkin, Yosi Keller, and Tal Hassner · 2019
Earlier work this paper cites.
Detecting deep-fake videos from appearance and behavior
Shruti Agarwal, Hany Farid, Tarek El-Gaaly, and Ser-Nam Lim · 2020
Earlier work this paper cites.
Evading deepfake-image detectors with white-and black-box attacks
Nicholas Carlini and Hany Farid · 2020
Earlier work this paper cites.
On the plausibility of virtual body animation features in virtual reality
Henrique Galvan Debarba, Sylvain Chagué, and Caecilia Charbonnier · 2020
Earlier work this paper cites.
Zhongang Cai, Mingyuan Zhang, Jiawei Ren, Chen Wei, Daxuan Ren, Zhengyu Lin, Haiyu Zhao, Lei Yang, Chen Change Loy, and Ziwei Liu · 2021
Earlier work this paper cites.
A meta-analysis of the uncanny valley’s independent and dependent variables
Alexander Diel, Sarah Weigelt, and Karl F Macdorman · 2021
Earlier work this paper cites.
Deepfake detection based on discrepancies between faces and their context
Yuval Nirkin, Lior Wolf, Yosi Keller, and Tal Hassner · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
LAION-400M: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki · 2021
Earlier work this paper cites.
How much can CLIP benefit vision-and-language tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer · 2021
Earlier work this paper cites.
Who are you (I really wanna know)? Detecting audio DeepFakes through vocal tract reconstruction
Logan Blue, Kevin Warren, Hadi Abdullah, Cassidy Gibson, Luis Vargas, Jessica O’Dell, Kevin Butler, and Patrick Traynor · 2022
Earlier work this paper cites.
Protecting world leaders against deep fakes using facial, gestural, and vocal mannerisms
Matyáš Boháček and Hany Farid · 2022
Earlier work this paper cites.
PaLI: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al · 2022
Cited alongside, same era.
Creating, using, misusing, and detecting deep fakes
Hany Farid · 2022
Cited alongside, same era.
Model attribution of face-swap deepfake videos
Shan Jia, Xin Li, and Siwei Lyu · 2022
Cited alongside, same era.
AI-synthesized faces are indistinguishable from real faces and more trustworthy
Sophie J Nightingale and Hany Farid · 2022
Cited alongside, same era.
Deepfake audio detection by speaker verification
Alessandro Pianese, Davide Cozzolino, Giovanni Poggi, and Luisa Verdoliva · 2022
Cited alongside, same era.
LAION-5B: An open large-scale dataset for training next generation image-text models
To authenticity, and beyond! Building safe and fair generative AI upon the three pillars of provenance
John Collomosse and Andy Parsons · 2024
Closest in time.
Raising the bar of AI-generated image detection with CLIP
Davide Cozzolino, Giovanni Poggi, Riccardo Corvi, Matthias Nießner, and Luisa Verdoliva · 2024
Closest in time.
Exposing lip-syncing deepfakes from mouth inconsistencies
Soumyya Kanti Datta, Shan Jia, and Siwei Lyu · 2024
Closest in time.
Exploring the adversarial robustness of CLIP for AI-generated image detection
Vincenzo De Rosa, Fabrizio Guillaro, Giovanni Poggi, Davide Cozzolino, and Luisa Verdoliva · 2024
Closest in time.
https://runwayml.com/research/introducing-gen-3-alpha , 2024
RunwayML Gen3 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Cited alongside, same era.
CLIP models are few-shot learners: Empirical studies on vqa and visual entailment
Haoyu Song, Li Dong, Wei-Nan Zhang, Ting Liu, and Furu Wei · 2022
Cited alongside, same era.
Single and multi-speaker cloned voice detection: From perceptual to learned features
Sarah Barrington, Romit Barua, Gautham Koorma, and Hany Farid · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al · 2023
Cited alongside, same era.
VideoPoet: A large language model for zero-shot video generation
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Rachel Hornung, Hartwig Adam, Hassan Akbari, Yair Alon, Vighnesh Birodkar, et al · 2023
Cited alongside, same era.
SMPL: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Cited alongside, same era.
Can ChatGPT detect deepfakes? A study of using multimodal large language models for media forensics
Shan Jia, Reilin Lyu, Kangran Zhao, Yize Chen, Zhiyuan Yan, Yan Ju, Chuanbo Hu, Xin Li, Baoyuan Wu, and Siwei Lyu · 2024
Closest in time.
CLIPping the deception: Adapting vision-language models for universal deepfake detection
Sohail Ahmed Khan and Duc-Tien Dang-Nguyen · 2024
Closest in time.
Jina CLIP: Your CLIP model is also your text retriever
Andreas Koukounas, Georgios Mastrapas, Michael Günther, Bo Wang, Scott Martens, Isabelle Mohr, Saba Sturua, Mohammad Kalim Akram, Joan Fontanals Martínez, Saahil Ognawala, et al · 2024
Closest in time.
AnimateDiff-Lightning: Cross-model diffusion distillation
Shanchuan Lin and Xiao Yang · 2024
Closest in time.
From audio to photoreal embodiment: Synthesizing humans in conversations
Evonne Ng, Javier Romero, Timur Bagautdinov, Shaojie Bai, Trevor Darrell, Angjoo Kanazawa, and Alexander Richard · 2024
Closest in time.
https://www.pexels.com , 2024
Pexels · 2024
Closest in time.
DeCLIP: Decoding CLIP representations for deepfake localization
Stefan Smeu, Elisabeta Oneata, and Dan Oneata · 2024
Closest in time.
https://openai.com/index/sora/ , 2024
Sora · 2024
Closest in time.
Beyond deepfake images: Detecting AI-generated videos
Danial Samadi Vahdati, Tai D Nguyen, Aref Azizpour, and Matthew C Stamm · 2024
Closest in time.
https://deepmind.google/technologies/veo , 2024
Veo · 2024
Closest in time.
CogVideoX: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al · 2024
Closest in time.
4Real: Towards photorealistic 4D scene generation via video diffusion models
Heng Yu, Chaoyang Wang, Peiye Zhuang, Willi Menapace, Aliaksandr Siarohin, Junli Cao, Laszlo A Jeni, Sergey Tulyakov, and Hsin-Ying Lee · 2024
Closest in time.