Fetching the paper…
Reading the bibliography…
Diffusion models have demonstrated superior performance in the field of portrait animation.
Styletalk: One-shot talking head generation with controllable speaking styles
Ma, Y.; Wang, S.; Hu, Z.; Fan, C.; Lv, T.; Ding, Y.; Deng, Z.; and Yu, X. 2023 · 1904
Earlier work this paper cites.
Mediapipe: A framework for building perception pipelines
Lugaresi, C.; Tang, J.; Nash, H.; McClanahan, C.; Uboweja, E.; Hays, M.; Zhang, F.; Chang, C.-L.; Yong, M. G.; Lee, J.; et al. 2019 · 1906
Earlier work this paper cites.
Efficient 3D Implicit Head Avatar with Mesh-anchored Hash Table Blendshapes
Bai, Z.; Tan, F.; Fanello, S.; Pandey, R.; Dou, M.; Liu, S.; Tan, P.; and Zhang, Y. 2024 · 1984
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Chung, J. S.; and Zisserman, A. 2016 · 2016
Earlier work this paper cites.
Learning a model of facial shape and expression from 4D scans
Li, T.; Bolkart, T.; Black, M. J.; Li, H.; and Romero, J. 2017 · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset
Nagrani, A.; Chung, J. S.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
Chung, J. S.; Nagrani, A.; and Zisserman, A. 2018 · 2018
Earlier work this paper cites.
Assessing empathy and managing emotions through interactions with an affective avatar
Johnson, E.; Hervás, R.; Gutiérrez López de la Franca, C.; Mondéjar, T.; Ochoa, S. F.; and Favela, J. 2018 · 2018
Earlier work this paper cites.
Nonlinear 3d face morphable model
Tran, L.; and Liu, X. 2018 · 2018
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Karras, T.; Laine, S.; and Aila, T. 2019 · 2019
Earlier work this paper cites.
Casual conversations: A dataset for measuring fairness in ai
Hazirbas, C.; Bitton, J.; Dolhansky, B.; Pan, J.; Gordo, A.; and Ferrer, C. C. 2021 · 2021
Earlier work this paper cites.
Pixel codec avatars
Ma, S.; Simon, T.; Saragih, J.; Wang, D.; Li, Y.; De La Torre, F.; and Sheikh, Y. 2021 · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Mildenhall, B.; Srinivasan, P. P.; Tancik, M.; Barron, J. T.; Ramamoorthi, R.; and Ng, R. 2021 · 2021
Earlier work this paper cites.
Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual Dataset
Zhang, Z.; Li, L.; Ding, Y.; and Fan, C. 2021 · 2021
Earlier work this paper cites.
Eamm: One-shot emotional talking face via audio-based emotion-aware motion model
Ji, X.; Zhou, H.; Wang, K.; Wu, Q.; Wu, W.; Xu, F.; and Cao, X. 2022 · 2022
Earlier work this paper cites.
Expressive talking head generation with granular audio-visual control
Liang, B.; Pan, Y.; Guo, Z.; Zhou, H.; Hong, Z.; Han, X.; Han, J.; Liu, J.; Ding, E.; and Wang, J. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Earlier work this paper cites.
Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis
Song, L.; Wu, W.; Fu, C.; Loy, C. C.; and He, R. 2022 · 2022
Earlier work this paper cites.
Vfhq: A high-quality dataset and benchmark for video face super-resolution
Xie, L.; Wang, X.; Zhang, H.; Dong, C.; and Shan, Y. 2022 · 2022
Earlier work this paper cites.
MagicDance: Realistic Human Dance Video Generation with Motions & Facial Expressions Transfer
Chang, D.; Shi, Y.; Gao, Q.; Fu, J.; Xu, H.; Song, G.; Yan, Q.; Yang, X.; and Soleymani, M. 2023 · 2023
Cited alongside, same era.
Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis
Chen, J.; Yu, J.; Ge, C.; Yao, L.; Xie, E.; Wu, Y.; Wang, Z.; Kwok, J.; Luo, P.; Lu, H.; et al. 2023 · 2023
Cited alongside, same era.
Efficient emotional adaptation for audio-driven talking-head generation
Gan, Y.; Yang, Z.; Yue, X.; Sun, L.; and Yang, Y. 2023 · 2023
Cited alongside, same era.
Stylesync: High-fidelity generalized and personalized lip sync in style-based generator
Guan, J.; Zhang, Z.; Zhou, H.; Hu, T.; Wang, K.; He, D.; Feng, H.; Liu, J.; Ding, E.; Liu, Z.; et al. 2023 · 2023
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
https://www.prnewswire.com/news-releases/deepbrain-ai-delivers-ai-avatar-to empower-people-with-disabilities-302026965.html
AI, D. 2024 · 2024
Closest in time.
NERF-AD: Neural Radiance Field With Attention-Based Disentanglement For Talking Face Synthesis
Bi, C.; Liu, X.; and Liu, Z. 2024 · 2024
Closest in time.
Cho, K.; Lee, J.; Yoon, H.; Hong, Y.; Ko, J.; Ahn, S.; and Kim, S. 2024 · 2024
Closest in time.
EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars
Drobyshev, N.; Casademunt, A. B.; Vougioukas, K.; Landgraf, Z.; Petridis, S.; and Pantic, M. 2024 · 2024
Closest in time.
LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control
Guo, J.; Zhang, D.; Liu, X.; Zhong, Z.; Zhang, Y.; Wan, P.; and Zhang, D. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guo, Y.; Yang, C.; Rao, A.; Liang, Z.; Wang, Y.; Qiao, Y.; Agrawala, M.; Lin, D.; and Dai, B. 2023 · 2023
Cited alongside, same era.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Hu, L.; Gao, X.; Zhang, P.; Sun, K.; Zhang, B.; and Bo, L. 2023 · 2023
Cited alongside, same era.
3D Gaussian Splatting for Real-Time Radiance Field Rendering
Kerbl, B.; Kopanas, G.; Leimkühler, T.; and Drettakis, G. 2023 · 2023
Cited alongside, same era.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
Khachatryan, L.; Movsisyan, A.; Tadevosyan, V.; Henschel, R.; Wang, Z.; Navasardyan, S.; and Shi, H. 2023 · 2023
Cited alongside, same era.
Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis
Li, J.; Zhang, J.; Bai, X.; Zhou, J.; and Gu, L. 2023 · 2023
Cited alongside, same era.
Moda: Mapping-once audio-driven portrait animation with dual attentions
Liu, Y.; Lin, L.; Yu, F.; Zhou, C.; and Li, Y. 2023 · 2023
Cited alongside, same era.
Videofusion: Decomposed diffusion models for high-quality video generation
Luo, Z.; Chen, D.; Zhang, Y.; Huang, Y.; Wang, L.; Shen, Y.; Zhao, D.; Zhou, J.; and Tan, T. 2023 · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; and Rombach, R. 2023 · 2023
Cited alongside, same era.
Closest in time.
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
Liu, T.; Chen, F.; Fan, S.; Du, C.; Chen, Q.; Chen, X.; and Yu, K. 2024 · 2024
Closest in time.
Synctalk: The devil is in the synchronization for talking head synthesis
Peng, Z.; Hu, W.; Shi, Y.; Zhu, X.; Zhang, X.; Zhao, H.; He, J.; Liu, H.; and Fan, Z. 2024 · 2024
Closest in time.
Adaptive Super Resolution for One-Shot Talking-Head Generation
Song, L.; Liu, P.; Yin, G.; and Xu, C. 2024 · 2024
Closest in time.
Diffused heads: Diffusion models beat gans on talking-face generation
Stypułkowski, M.; Vougioukas, K.; He, S.; Zieba, M.; Petridis, S.; and Pantic, M. 2024 · 2024
Closest in time.
Audio-driven High-resolution Seamless Talking Head Video Editing via StyleGAN
Su, J.; Liu, K.; Chen, L.; Yao, J.; Liu, Q.; and Lv, D. 2024 · 2024
Closest in time.
DT-NeRF: Decomposed Triplane-Hash Neural Radiance Fields For High-Fidelity Talking Portrait Synthesis
Su, Y.; Wang, S.; and Wang, H. 2024 · 2024
Closest in time.
MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset
Sung-Bin, K.; Chae-Yeon, L.; Son, G.; Hyun-Bin, O.; Ju, J.; Nam, S.; and Oh, T.-H. 2024 · 2024
Closest in time.
Tian, L.; Wang, Q.; Zhang, B.; and Bo, L. 2024 · 2024
Closest in time.
Aniportrait: Audio-driven synthesis of photorealistic portrait animation
Wei, H.; Yang, Z.; and Wang, Z. 2024 · 2024
Closest in time.
X-Portrait: Expressive Portrait Animation with Hierarchical Motion Attention
Xie, Y.; Xu, H.; Song, G.; Wang, C.; Shi, Y.; and Luo, L. 2024 · 2024
Closest in time.
MegActor: Harness the Power of Raw Video for Vivid Portrait Animation
Yang, S.; Li, H.; Wu, J.; Jing, M.; Li, L.; Ji, R.; Liang, J.; and Fan, H. 2024 · 2024
Closest in time.
Yu, R.; He, T.; Zeng, A.; Wang, Y.; Guo, J.; Tan, X.; Liu, C.; Chen, J.; and Bian, J. 2024 · 2024
Closest in time.
Champ: Controllable and consistent human image animation with 3d parametric guidance
Zhu, S.; Chen, J. L.; Dai, Z.; Xu, Y.; Cao, X.; Yao, Y.; Zhu, H.; and Zhu, S. 2024 · 2024
Closest in time.
Learn2Talk: 3D Talking Face Learns from 2D Talking Face
Zhuang, Y.; Cheng, B.; Cheng, Y.; Jin, Y.; Liu, R.; Li, C.; Cheng, X.; Liao, J.; and Lin, J. 2024 · 2024
Closest in time.