Fetching the paper…
Reading the bibliography…
In this paper, we abstract the process of people hearing speech, extracting meaningful cues, and creating various dynamically audio-consistent talking faces, termed Listening and Imagining, into the task of high-fidelity diverse talking faces generation from a single audio.
Facial action coding system
Paul Ekman and Wallace V Friesen · 1978
Earlier work this paper cites.
The voice-recognition accuracy of blind listeners
Ray Bull, Harriet Rathborn, and Brian R Clifford · 1983
Earlier work this paper cites.
Role of bone conduction in the self-perception of speech
Dieter Maurer and Theodor Landis · 1990
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Understanding the mechanisms of economic development
Angus Deaton · 2010
Earlier work this paper cites.
The handbook of phonetic sciences
William J Hardcastle, John Laver, and Fiona E Gibbon · 2012
Earlier work this paper cites.
Speaker age classification and regression using i-vectors
Joanna Grzybowska and Stanislaw Kacprzak · 2016
Earlier work this paper cites.
The relationship of voice onset time and voice offset time to physical age
Rita Singh, Joseph Keshet, Deniz Gencaga, and Bhiksha Raj · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2017
Earlier work this paper cites.
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Learning utterance-level representations for speech emotion and age/gender recognition using deep neural networks
Zhong-Qiu Wang and Ivan Tashev · 2017
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Speaker recognition from raw waveform with sincnet
Mirco Ravanelli and Yoshua Bengio · 2018
Earlier work this paper cites.
Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set
Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong · 2019
Earlier work this paper cites.
Emotional expression: Advances in basic emotion theory
Dacher Keltner, Disa Sauter, Jessica Tracy, and Alan Cowen · 2019
Earlier work this paper cites.
Improving transformer-based speech recognition systems with compressed structure and speech attributes augmentation
Sheng Li, Raj Dabre, Xugang Lu, Peng Shen, Tatsuya Kawahara, and Hisashi Kawai · 2019
Earlier work this paper cites.
frame attention networks for facial expression recognition in videos
Debin Meng, Xiaojiang Peng, Kai Wang, and Yu Qiao · 2019
Earlier work this paper cites.
Speech2face: Learning the face behind a voice
Tae-Hyun Oh, Tali Dekel, Changil Kim, Inbar Mosseri, William T Freeman, Michael Rubinstein, and Wojciech Matusik · 2019
Earlier work this paper cites.
Face reconstruction from voice using generative adversarial networks
Yandong Wen, Bhiksha Raj, and Rita Singh · 2019
Earlier work this paper cites.
Attention-augmented end-to-end multi-task learning for emotion prediction from speech
Zixing Zhang, Bingwen Wu, and Björn Schuller · 2019
Earlier work this paper cites.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Cited alongside, same era.
From inference to generation: End-to-end fully self-supervised generation of human face from speech
Hyeong-Seok Choi, Changdae Park, and Kyogu Lee · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Cited alongside, same era.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Cited alongside, same era.
Styleheat: One-shot high-resolution editable talking face generation via pretrained stylegan
Fei Yin, Yong Zhang, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang, Qingyan Bai, Baoyuan Wu, Jue Wang, and Yujiu Yang · 2022
Later among the works it cites.
Celebv-hq: A large-scale video facial attributes dataset
Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy · 2022
Later among the works it cites.
Control-a-video: Controllable text-to-video generation with diffusion models
Weifeng Chen, Jie Wu, Pan Xie, Hefeng Wu, Jiashi Li, Xin Xia, Xuefeng Xiao, and Liang Lin · 2023
Later among the works it cites.
Dae-talker: High fidelity speech-driven talking face generation with diffusion autoencoder
Chenpeng Du, Qi Chen, Tianyu He, Xu Tan, Xie Chen, Kai Yu, Sheng Zhao, and Jiang Bian · 2023
Later among the works it cites.
Structure and content-guided video synthesis with diffusion models
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mead: A large-scale audio-visual dataset for emotional talking-face generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy · 2020
Cited alongside, same era.
Audio-driven emotional video portraits
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu · 2021
Cited alongside, same era.
Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation
Kimmo Karkkainen and Jungseock Joo · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Pirenderer: Controllable portrait image generation via semantic neural rendering
Yurui Ren, Ge Li, Yuanqi Chen, Thomas H Li, and Shan Liu · 2021
Cited alongside, same era.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan · 2021
Cited alongside, same era.
Later among the works it cites.
High-fidelity and freely controllable talking head video generation
Yue Gao, Yuan Zhou, Jinglu Wang, Xiao Li, Xiang Ming, and Yan Lu · 2023
Later among the works it cites.
Composer: Creative and controllable image synthesis with composable conditions
Lianghua Huang, Di Chen, Yu Liu, Yujun Shen, Deli Zhao, and Jingren Zhou · 2023
Later among the works it cites.
Gligen: Open-set grounded text-to-image generation
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee · 2023
Later among the works it cites.
Follow your pose: Pose-guided text-to-video generation using pose-free videos
Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Ying Shan, Xiu Li, and Qifeng Chen · 2023
Later among the works it cites.
Chong Mou, Xintao Wang, Liangbin Xie, Jian Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie · 2023
Later among the works it cites.
Dpe: Disentanglement of pose and expression for general video portrait editing
Youxin Pang, Yong Zhang, Weize Quan, Yanbo Fan, Xiaodong Cun, Ying Shan, and Dong-ming Yan · 2023
Later among the works it cites.
Emotalk: Speech-driven emotional disentanglement for 3d face animation
Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xiangyu Zhu, Jun He, Hongyan Liu, and Zhaoxin Fan · 2023
Later among the works it cites.
Fatezero: Fusing attentions for zero-shot text-based video editing
Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen · 2023
Later among the works it cites.
Difftalk: Crafting diffusion models for generalized talking head synthesis
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu · 2023
Later among the works it cites.
Instantbooth: Personalized text-to-image generation without test-time finetuning
Jing Shi, Wei Xiong, Zhe Lin, and Hyun Joon Jung · 2023
Later among the works it cites.
Diffused heads: Diffusion models beat gans on talking-face generation
Michal Stypulkowski, Konstantinos Vougioukas, Sen He, Maciej Zikeba, Stavros Petridis, and Maja Pantic · 2023
Later among the works it cites.
Emmn: Emotional motion memory network for audio-driven emotional talking face generation
Shuai Tan, Bin Ji, and Ye Pan · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou · 2023
Later among the works it cites.
Simda: Simple diffusion adapter for efficient video generation
Zhen Xing, Qi Dai, Han Hu, Zuxuan Wu, and Yu-Gang Jiang · 2023
Later among the works it cites.
Inpaint anything: Segment anything meets image inpainting
Tao Yu, Runseng Feng, Ruoyu Feng, Jinming Liu, Xin Jin, Wenjun Zeng, and Zhibo Chen · 2023
Later among the works it cites.
Identity-preserving talking face generation with landmark and appearance priors
Weizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei, Gangming Zhao, Liang Lin, and Guanbin Li · 2023
Later among the works it cites.