Fetching the paper…
Reading the bibliography…
Human-centric generative models are becoming increasingly popular, giving rise to various innovative tools and applications, such as talking face videos conditioned on text or audio prompts.
No-reference image quality assessment in the spatial domain
Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik · 2012
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly · 2018
Earlier work this paper cites.
First order motion model for image animation
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe · 2019
Earlier work this paper cites.
Paddleocr, awesome multilingual ocr toolkits based on paddlepaddle
PaddlePaddle Authors · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Earlier work this paper cites.
Towards measuring fairness in ai: the casual conversations dataset
Caner Hazirbas, Joanna Bitton, Brian Dolhansky, Jacqueline Pan, Albert Gordo, and Cristian Canton Ferrer · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Earlier work this paper cites.
Write-a-speaker: Text-based emotional and rhythmic talking-head generation
Lincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding, Yixing Zheng, Xin Yu, and Changjie Fan · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Earlier work this paper cites.
Motion representations for articulated animation
Aliaksandr Siarohin, Oliver J Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Earlier work this paper cites.
One-shot free-view neural talking-head synthesis for video conferencing
Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu · 2021
Earlier work this paper cites.
Deconfounded video moment retrieval with causal intervention
Xun Yang, Fuli Feng, Wei Ji, Meng Wang, and Tat-Seng Chua · 2021
Earlier work this paper cites.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan · 2021
Earlier work this paper cites.
Grainspace: A large-scale dataset for fine-grained and domain-adaptive recognition of cereal grains
Lei Fan, Yiwen Ding, Dongdong Fan, Donglin Di, Maurice Pagnucco, and Yang Song · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Video moment retrieval with cross-modal neural architecture search
Xun Yang, Shanshan Wang, Jian Dong, Jianfeng Dong, Meng Wang, and Tat-Seng Chua · 2022
Cited alongside, same era.
Text2video: Text-driven talking-head video synthesis with personalized phoneme-pose dictionary
Sibo Zhang, Jiahong Yuan, Miao Liao, and Liangjun Zhang · 2022
Cited alongside, same era.
Thin-plate spline motion model for image animation
Jian Zhao and Hui Zhang · 2022
Cited alongside, same era.
Towards robust blind face restoration with codebook lookup transformer
Shangchen Zhou, Kelvin C.K. Chan, Chongyi Li, and Chen Change Loy · 2022
Cited alongside, same era.
Celebv-hq: A large-scale video facial attributes dataset
Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy · 2022
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Emoportraits: Emotion-enhanced multimodal one-shot head avatars
Nikita Drobyshev, Antoni Bigata Casademunt, Konstantinos Vougioukas, Zoe Landgraf, Stavros Petridis, and Maja Pantic · 2024
Closest in time.
One-shot pose-driving face animation platform
He Feng, Donglin Di, Yongjia Ma, Wei Chen, and Tonghua Su · 2024
Closest in time.
Psycollm: Enhancing llm for psychological understanding and evaluation
Jinpeng Hu, Tengteng Dong, Luo Gang, Hui Ma, Peng Zou, Xiao Sun, Dan Guo, Xun Yang, and Meng Wang · 2024
Closest in time.
Faces that speak: Jointly synthesising talking face and speech from text
Youngjoon Jang, Ji-Hoon Kim, Junseok Ahn, et al · 2024
Closest in time.
Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization
Yang Jin, Zhicheng Sun, Kun Xu, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andreas Blattmann, Tim Dockhorn, and Sumith andothers Kulal · 2023
Cited alongside, same era.
An annotated grain kernel image database for visual quality inspection
Lei Fan, Yiwen Ding, Dongdong Fan, Yong Wu, Hongxia Chu, Maurice Pagnucco, and Yang Song · 2023
Cited alongside, same era.
Efficient emotional adaptation for audio-driven talking-head generation
Yuan Gan, Zongxin Yang, Xihang Yue, Lingyun Sun, and Yi Yang · 2023
Cited alongside, same era.
Implicit identity representation conditioned memory compensation network for talking head video generation
Fa-Ting Hong and Dan Xu · 2023
Cited alongside, same era.
Moda: Mapping-once audio-driven portrait animation with dual attentions
Yunfei Liu, Lijian Lin, Fei Yu, Changyin Zhou, and Yu Li · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Cited alongside, same era.
The casual conversations v2 dataset
Bilal Porgali, Vítor Albiero, Jordan Ryda, Cristian Canton Ferrer, and Caner Hazirbas · 2023
Cited alongside, same era.
Latentsync: Audio conditioned latent diffusion models for lip sync
Chunyu Li, Chao Zhang, Weikai Xu, et al · 2024
Closest in time.
Emotional conversation: Empowering talking faces with cohesive expression, gaze and pose generation
Jiadong Liang and Feng Lu · 2024
Closest in time.
Facexformer: A unified transformer for facial analysis
Kartik Narayan, Vibashan VS, Rama Chellappa, and Vishal M Patel · 2024
Closest in time.
Consisti2v: Enhancing visual consistency for image-to-video generation
Weiming Ren, Harry Yang, Ge Zhang, et al · 2024
Closest in time.
Diffused heads: Diffusion models beat gans on talking-face generation
Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zikeba, Stavros Petridis, and Maja Pantic · 2024
Closest in time.
Aniportrait: Audio-driven synthesis of photorealistic portrait animation
Huawei Wei, Zejun Yang, and Zhisheng Wang · 2024
Closest in time.
X-portrait: Expressive portrait animation with hierarchical motion attention
You Xie, Hongyi Xu, Guoxian Song, Chao Wang, Yichun Shi, and Linjie Luo · 2024
Closest in time.
Real3d-portrait: One-shot realistic 3d talking portrait synthesis
Zhenhui Ye, Tianyun Zhong, Yi Ren, et al · 2024
Closest in time.
Open-sora: Democratizing efficient video production for all, 2024
Zangwei Zheng, Xiangyu Peng, Tianji Yang, et al · 2024
Closest in time.
VBench: Comprehensive benchmark suite for video generative models
Ziqi Huang, Yinan He, Jiashuo Yu, et al · 2024
Closest in time.
Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer
Jiahao Cui, Hui Li, Yun Zhan, Hanlin Shang, Kaihui Cheng, Yuqi Ma, Shan Mu, Hang Zhou, Jingdong Wang, and Siyu Zhu · 2025
Closest in time.
Hyper-3dg: Text-to-3d gaussian generation via hypergraph
Donglin Di, Jiahui Yang, Chaofan Luo, Zhou Xue, Wei Chen, Xun Yang, and Yue Gao · 2025
Closest in time.
Manta: A large-scale multi-view and visual-text anomaly detection dataset for tiny objects
Lei Fan, Dongdong Fan, Zhiguang Hu, et al · 2025
Closest in time.
Adams bashforth moulton solver for inversion and editing in rectified flow
Yongjia Ma, Donglin Di, Xuan Liu, Xiaokai Chen, Lei Fan, Wei Chen, and Tonghua Su · 2025
Closest in time.
Precise localization of memories: A fine-grained neuron-level knowledge editing technique for llms
Haowen Pan, Xiaozhi Wang, Yixin Cao, Zenglin Shi, Xun Yang, Juanzi Li, and Meng Wang · 2025
Closest in time.
Egotextvqa: Towards egocentric scene-text aware video question answering
Sheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li, Xun Yang, Dan Guo, Meng Wang, Tat-Seng Chua, and Angela Yao · 2025
Closest in time.