Fetching the paper…
Reading the bibliography…
Imagine having a conversation with a socially intelligent agent.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
Out of time: automated lip sync in the wild
J. S. Chung and A. Zisserman · 2016
Earlier work this paper cites.
Neural discrete representation learning, 2018
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Mediapipe: A framework for perceiving and processing reality
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Yong, Juhyun Lee, Wan-Teh Chang, Wei Hua, Manfred Georg, and Matthias Grundmann · 2019
Earlier work this paper cites.
Talknet: Fully-convolutional non-autoregressive speech synthesis model, 2020
Stanislav Beliaev, Yurii Rebryk, and Boris Ginsburg · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Analyzing and improving the image quality of StyleGAN
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2020
Earlier work this paper cites.
A lip sync expert is all you need for speech to lip generation in the wild
K R Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, and C.V. Jawahar · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Earlier work this paper cites.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Earlier work this paper cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Earlier work this paper cites.
Pirenderer: Controllable portrait image generation via semantic neural rendering, 2021
Yurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li, and Shan Liu · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Earlier work this paper cites.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan · 2021
Earlier work this paper cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Cited alongside, same era.
Megaportraits: One-shot megapixel neural head avatars
Nikita Drobyshev, Jenya Chelishev, Taras Khakhulin, Aleksei Ivakhnenko, Victor Lempitsky, and Egor Zakharov · 2022
Cited alongside, same era.
Visual speech-aware perceptual 3d facial expression reconstruction from videos
Panagiotis P. Filntisis, George Retsinas, Foivos Paraperas-Papantoniou, Athanasios Katsamanis, Anastasios Roussos, and Petros Maragos · 2022
Cited alongside, same era.
Perceptual conversational head generation with regularized driver and enhanced renderer
Ailin Huang, Zhewei Huang, and Shuchang Zhou · 2022
Cited alongside, same era.
Learning to listen: Modeling non-deterministic dyadic facial motion
Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, and Shiry Ginosar · 2022
Cited alongside, same era.
Dialoguenerf: Towards realistic avatar face-to-face conversation video generation, 2023
Yichao Yan, Zanwei Zhou, Zi Wang, Jingnan Gao, and Xiaokang Yang · 2023
Later among the works it cites.
Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation, 2023
Shengkui Zhao, Yukun Ma, Chongjia Ni, Chong Zhang, Hao Wang, Trung Hieu Nguyen, Kun Zhou, Jiaqi Yip, Dianwen Ng, and Bin Ma · 2023
Later among the works it cites.
Interactive conversational head generation
Mohan Zhou, Yalong Bai, Wei Zhang, Ting Yao, and Tiejun Zhao · 2023
Later among the works it cites.
Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions
Zhiyuan Chen, Jiajiong Cao, Zhiquan Chen, Yuming Li, and Chenguang Ma · 2024
Closest in time.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Real-time neural radiance talking portrait synthesis via audio-spatial decomposition
Jiaxiang Tang, Kaisiyuan Wang, Hang Zhou, Xiaokang Chen, Dongliang He, Tianshu Hu, Jingtuo Liu, Gang Zeng, and Jingdong Wang · 2022
Cited alongside, same era.
Latent image animator: Learning to animate images via latent space navigation
Yaohui Wang, Di Yang, Francois Bremond, and Antitza Dantcheva · 2022
Cited alongside, same era.
Semantic-aware responsive listener head synthesis
Wei Zhao, Peng Xiao, Rongju Zhang, Yijun Wang, and Jianxin Lin · 2022
Cited alongside, same era.
Vico-x: Multimodal conversation dataset
Mohan Zhou, Yalong Bai, Wei Zhang, Ting Yao, Tiejun Zhao, and Tao Mei · 2022
Cited alongside, same era.
Affective faces for goal-driven dyadic communication
Scott Geng, Revant Teotia, Purva Tendulkar, Sachit Menon, and Carl Vondrick · 2023
Cited alongside, same era.
Stylesync: High-fidelity generalized and personalized lip sync in style-based generator
Jiazhi Guan, Zhanwang Zhang, Hang Zhou, Tianshu Hu, Kaisiyuan Wang, Dongliang He, Haocheng Feng, Jingtuo Liu, Errui Ding, Ziwei Liu, et al · 2023
Cited alongside, same era.
Mfr-net: Multi-faceted responsive listening head generation via denoising diffusion model
Jin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, and Jizhong Han · 2023
Cited alongside, same era.
Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai · 2024
Closest in time.
Interact: Capture and modelling of realistic, expressive and interactive activities between two persons in daily scenarios
Yinghao Huang, Leo Ho, Dafei Qin, Mingyi Shi, and Taku Komura · 2024
Closest in time.
Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation
Yue Ma, Hongyu Liu, Hongfa Wang, Heng Pan, Yingqing He, Junkun Yuan, Ailing Zeng, Chengfei Cai, Heung-Yeung Shum, Wei Liu, et al · 2024
Closest in time.
From audio to photoreal embodiment: Synthesizing humans in conversations
Evonne Ng, Javier Romero, Timur Bagautdinov, Shaojie Bai, Trevor Darrell, Angjoo Kanazawa, and Alexander Richard · 2024
Closest in time.
Let’s go real talk: Spoken dialogue model for face-to-face conversation
Se Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim, Joanna Hong, Jeong Hun Yeo, and Yong Man Ro · 2024
Closest in time.
React 2024: the second multiple appropriate facial reaction generation challenge, 2024
Siyang Song, Micol Spitale, Cheng Luo, Cristina Palmero, German Barquero, Hengde Zhu, Sergio Escalera, Michel Valstar, Tobias Baur, Fabien Ringeval, Elisabeth Andre, and Hatice Gunes · 2024
Closest in time.
Diffused heads: Diffusion models beat gans on talking-face generation
Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zięba, Stavros Petridis, and Maja Pantic · 2024
Closest in time.
Linrui Tian, Qi Wang, Bang Zhang, and Liefeng Bo · 2024
Closest in time.
Dyadic interaction modeling for social behavior generation
Minh Tran, Di Chang, Maksim Siniukov, and Mohammad Soleymani · 2024
Closest in time.
V-express: Conditional dropout for progressive training of portrait video generation
Cong Wang, Kuan Tian, Jun Zhang, Yonghang Guan, Feng Luo, Fei Shen, Zhiwei Jiang, Qing Gu, Xiao Han, and Wei Yang · 2024
Closest in time.
Personatalk: Bring attention to your persona in visual dubbing
Longhao Zhang, Shuang Liang, Zhipeng Ge, and Tianshu Hu · 2024
Closest in time.