Fetching the paper…
Reading the bibliography…
Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation.
A database of German emotional speech
Felix Burkhardt, Astrid Paeschke, Miriam Rolfes, Walter F Sendlmeier, Benjamin Weiss, et al · 2005
Earlier work this paper cites.
The enterface’05 audio-visual emotion database
Olivier Martin, Irene Kotsia, Benoit Macq, and Ioannis Pitas · 2006
Earlier work this paper cites.
Toronto emotional speech set (tess)-younger talker_happy
Kate Dupuis and M Kathleen Pichora-Fuller · 2010
Earlier work this paper cites.
The PAVOQUE corpus as a resource for analysis and synthesis of expressive speech
Ingmar Steiner, Marc Schröder, and Annette Klepp · 2013
Earlier work this paper cites.
Surrey audio-visual expressed emotion (savee) database
Philip Jackson and SJUoSG Haq · 2014
Earlier work this paper cites.
Polish emotional natural speech database
Dorota Kaminska, Tomasz Sapinski, and Adam Pelikant · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Out of time: Automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2017
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Synthesizing obama: Learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Earlier work this paper cites.
The emotional voices database: Towards controlling the emotion dimension in voice generation systems
Adaeze Adigwe, Noé Tits, Kevin El Haddad, Sarah Ostadabbas, and Thierry Dutoit · 2018
Earlier work this paper cites.
Lip movements generation at a glance
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
A Canadian French emotional speech dataset
Philippe Gournay, Olivier Lahaie, and Roch Lefebvre · 2018
Earlier work this paper cites.
An open source emotional speech corpus for human robot interaction applications
Jesin James, Li Tian, and Catherine Watson · 2018
Earlier work this paper cites.
Cross lingual speech emotion recognition: Urdu vs. Western languages
Siddique Latif, Adnan Qayyum, Muhammad Usman, and Junaid Qadir · 2018
Earlier work this paper cites.
The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english
Steven R Livingstone and Frank A Russo · 2018
Earlier work this paper cites.
Speech emotion recognition for performance interaction
Nikolaos Vryzas, Rigas Kotsakis, Aikaterini Liatsou, Charalampos A Dimoulas, and George Kalliris · 2018
Earlier work this paper cites.
The MTG-Jamendo dataset for automatic music tagging
Dmitry Bogdanov, Minz Won, Philip Tovstogan, Alastair Porter, and Xavier Serra · 2019
Earlier work this paper cites.
FVD: A new metric for video generation
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphaël Marinier, Marcin Michalski, and Sylvain Gelly · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Turkish emotion voice database (TurEV-DB)
Salih Firat Canpolat, Zuhal Ormanoğlu, and Deniz Zeyrek · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
French emotional speech database-oréau, 2020
Leila KERKENI, Catherine CLEDER, Youssef Serrestou, and Y Raood · 2020
Cited alongside, same era.
ASVP-ESD: A dataset and its benchmark for emotion recognition using both speech and non-speech utterances
Dejoli Landry, Qianhua He, Haikang Yan, and Yanxiong Li · 2020
Cited alongside, same era.
Voxceleb: Large-scale speaker verification in the wild
Arsha Nagrani, Joon Son Chung, Weidi Xie, and Andrew Zisserman · 2020
Cited alongside, same era.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Speech emotion recognition in Italian using wav2vec 2
Fabio Catania · 2023
Later among the works it cites.
AnimateDiff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai · 2023
Later among the works it cites.
GAIA: Zero-shot talking avatar generation
Tianyu He, Junliang Guo, Runyi Yu, Yuchi Wang, Kaikai An, Leyi Li, Xu Tan, Chunyu Wang, Han Hu, HsiangTao Wu, et al · 2023
Later among the works it cites.
Dreamtalk: When expressive talking head generation meets diffusion probabilistic models
Yifeng Ma, Shiwei Zhang, Jiayu Wang, Xiang Wang, Yingya Zhang, and Zhidong Deng · 2023
Later among the works it cites.
Kari Ali Noriy, Xiaosong Yang, and Jian Jun Zhang · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Blindly assess image quality in the wild guided by a self-adaptive hyper network
Shaolin Su, Qingsen Yan, Yu Zhu, Cheng Zhang, Xin Ge, Jinqiu Sun, and Yanning Zhang · 2020
Cited alongside, same era.
MEAD: A large-scale audio-visual dataset for emotional talking-face generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy · 2020
Cited alongside, same era.
Makelttalk: Speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Cited alongside, same era.
The Mexican emotional speech database (mesd): Elaboration and assessment based on machine learning
Mathilde M Duville, Luz M Alonso-Valerdi, and David I Ibarra-Zarate · 2021
Cited alongside, same era.
Sust bangla emotional speech corpus (subesco): An audio-only emotional speech corpus for bangla
Sadia Sultana, M Shahidur Rahman, M Reza Selim, and M Zafar Iqbal · 2021
Cited alongside, same era.
Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li · 2021
Cited alongside, same era.
Later among the works it cites.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Later among the works it cites.
SDXL: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach · 2023
Later among the works it cites.
A new amharic speech emotion dataset and classification benchmark
Ephrem Afele Retta, Eiad Almekhlafi, Richard Sutcliffe, Mustafa Mhamed, Haider Ali, and Jun Feng · 2023
Later among the works it cites.
Vividtalk: One-shot audio-driven talking head generation based on 3D hybrid prior
Xusen Sun, Longhao Zhang, Hao Zhu, Peng Zhang, Bang Zhang, Xinya Ji, Kangneng Zhou, Daiheng Gao, Liefeng Bo, and Xun Cao · 2023
Later among the works it cites.
DynamiCrafter: Animating open-domain images with video diffusion priors
Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Xintao Wang, Tien-Tsin Wong, and Ying Shan · 2023
Later among the works it cites.
Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions
Zhiyuan Chen, Jiajiong Cao, Zhiquan Chen, Yuming Li, and Chenguang Ma · 2024
Closest in time.
Hallo2: Long-duration and high-resolution audio-driven portrait image animation
Jiahao Cui, Hui Li, Yao Yao, Hao Zhu, Hanlin Shang, Kaihui Cheng, Hang Zhou, Siyu Zhu, and Jingdong Wang · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Closest in time.
Loopy: Taming audio-driven portrait avatar with long-term motion dependency
Jianwen Jiang, Chao Liang, Jiaqi Yang, Gaojie Lin, Tianyun Zhong, and Yanbo Zheng · 2024
Closest in time.
Transnet v2: An effective deep network architecture for fast shot transition detection
Tomás Soucek and Jakub Lokoc · 2024
Closest in time.
Diffused heads: Diffusion models beat gans on talking-face generation
Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zięba, Stavros Petridis, and Maja Pantic · 2024
Closest in time.
MultiTalk: Enhancing 3D talking head generation across languages with multilingual video dataset
Kim Sung-Bin, Lee Chae-Yeon, Gihun Son, Oh Hyun-Bin, Janghoon Ju, Suekyeong Nam, and Tae-Hyun Oh · 2024
Closest in time.
Flowvqtalker: High-quality emotional talking face generation through normalizing flow and quantization
Shuai Tan, Bin Ji, and Ye Pan · 2024
Closest in time.
EMO: Emote portrait alive-generating expressive portrait videos with audio2video diffusion model under weak conditions
Linrui Tian, Qi Wang, Bang Zhang, and Liefeng Bo · 2024
Closest in time.
V-Express: Conditional dropout for progressive training of portrait video generation
Cong Wang, Kuan Tian, Jun Zhang, Yonghang Guan, Feng Luo, Fei Shen, Zhiwei Jiang, Qing Gu, Xiao Han, and Wei Yang · 2024
Closest in time.
AniPortrait: Audio-driven synthesis of photorealistic portrait animation
Huawei Wei, Zejun Yang, and Zhisheng Wang · 2024
Closest in time.