Fetching the paper…
Reading the bibliography…
The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Wei, J.; and Zou, K. 2019 · 1901
Earlier work this paper cites.
Facial action coding system
Ekman, P.; and Friesen, W. V. 1978 · 1978
Earlier work this paper cites.
Visualizing data using t-SNE
Van der Maaten, L.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
CREMA-D: Crowd-Sourced Emotional Multimodal Actors Dataset
Cao, H.; Cooper, D. G.; Keutmann, M. K.; Gur, R. C.; Nenkova, A.; and Verma, R. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Karras, T.; Aila, T.; Laine, S.; Herva, A.; and Lehtinen, J. 2017 · 2017
Earlier work this paper cites.
Emotion Intensities in Tweets
Mohammad, S. M.; and Bravo-Marquez, F. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Capture, Learning, and Synthesis of 3D Speaking Styles
Cudeiro, D.; Bolkart, T.; Laidlaw, C.; Ranjan, A.; and Black, M. J. 2019 · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
Speech-driven expressive talking lips with conditional sequential generative adversarial networks
Sadoughi, N.; and Busso, C. 2019 · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A.; Zhou, Y.; Mohamed, A.; and Auli, M. 2020 · 2020
Earlier work this paper cites.
SciPy 1.0: fundamental algorithms for scientific computing in Python
Virtanen, P.; Gommers, R.; Oliphant, T. E.; Haberland, M.; Reddy, T.; Cournapeau, D.; Burovski, E.; Peterson, P.; Weckesser, W.; Bright, J.; et al. 2020 · 2020
Cited alongside, same era.
Mead: A large-scale audio-visual dataset for emotional talking-face generation
Wang, K.; Wu, Q.; Song, L.; Yang, Z.; Wu, W.; Qian, C.; He, R.; Qiao, Y.; and Loy, C. C. 2020 · 2020
Cited alongside, same era.
Audio-driven emotional video portraits
Ji, X.; Zhou, H.; Wang, K.; Wu, W.; Loy, C. C.; Cao, X.; and Xu, F. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement
Richard, A.; Zollhoefer, M.; Wen, Y.; la Torre, F. D.; and Sheikh, Y. 2021 · 2021
Cited alongside, same era.
Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis
Liu, H.; Zhu, Z.; Iwamoto, N.; Peng, Y.; Li, Z.; Zhou, Y.; Bozkurt, E.; and Zheng, B. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents.(arXiv preprint)(2022)
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2023 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
Emotion-controllable generalized talking face generation
Sinha, S.; Biswas, S.; Yadav, R.; and Bhowmick, B. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schaldenbrand, P.; Liu, Z.; and Oh, J. 2021 · 2021
Cited alongside, same era.
Imitating arbitrary talking style for realistic audio-driven talking face synthesis
Wu, H.; Jia, J.; Wang, H.; Dou, Y.; Duan, C.; and Deng, Q. 2021 · 2021
Cited alongside, same era.
Transformer-s2a: Robust and efficient speech-to-animation
Chen, L.; Wu, Z.; Ling, J.; Li, R.; Tan, X.; and Zhao, S. 2022 · 2022
Cited alongside, same era.
FaceFormer: Speech-Driven 3D Facial Animation with Transformers
Fan, Y.; Lin, Z.; Saito, J.; Wang, W.; and Komura, T. 2021 · 2022
Cited alongside, same era.
Perspective reconstruction of human faces by joint mesh and landmark regression
Guo, J.; Yu, J.; Lattas, A.; and Deng, J. 2022 · 2022
Cited alongside, same era.
Eamm: One-shot emotional talking face via audio-based emotion-aware motion model
Ji, X.; Zhou, H.; Wang, K.; Wu, Q.; Wu, W.; Xu, F.; and Cao, X. 2022 · 2022
Cited alongside, same era.
Abaw: Valence-arousal estimation, expression recognition, action unit detection & multi-task learning challenges
Kollias, D. 2022 · 2022
Cited alongside, same era.
Motionclip: Exposing human motion generation to clip space
Tevet, G.; Gordon, B.; Hertz, A.; Bermano, A. H.; and Cohen-Or, D. 2022 · 2022
Later among the works it cites.
HiFace: High-Fidelity 3D Face Reconstruction by Learning Static and Dynamic Details
Chai, Z.; Zhang, T.; He, T.; Tan, X.; Baltrusaitis, T.; Wu, H.; Li, R.; Zhao, S.; Yuan, C.; and Bian, J. 2023 · 2023
Closest in time.
Emotional Speech-Driven Animation with Content-Emotion Disentanglement
Daněček, R.; Chhatre, K.; Tripathi, S.; Wen, Y.; Black, M. J.; and Bolkart, T. 2023 · 2023
Closest in time.
A Hierarchical Representation Network for Accurate and Detailed Face Reconstruction from In-The-Wild Images
Lei, B.; Ren, J.; Feng, M.; Cui, M.; and Xie, X. 2023 · 2023
Closest in time.
Styletalk: One-shot talking head generation with controllable speaking styles
Ma, Y.; Wang, S.; Hu, Z.; Fan, C.; Lv, T.; Ding, Y.; Deng, Z.; and Yu, X. 2023 · 2023
Closest in time.
Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head Synthesis
Wang, D.; Deng, Y.; Yin, Z.; Shum, H.-Y.; and Wang, B. 2023 · 2023
Closest in time.
Codetalker: Speech-driven 3d facial animation with discrete motion prior
Xing, J.; Xia, M.; Zhang, Y.; Cun, X.; Wang, J.; and Wong, T.-T. 2023 · 2023
Closest in time.
Human-computer interaction system: A survey of talking-head generation
Zhen, R.; Song, W.; He, Q.; Cao, J.; Shi, L.; and Luo, J. 2023 · 2023
Closest in time.