Fetching the paper…
Reading the bibliography…
Recent advances in conditional diffusion models have shown promise for generating realistic TalkingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A. C., Sheikh, H. R., and Simoncelli, E. P · 2004
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
Mirza, M. and Osindero, S · 2014
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Chung, J. S. and Zisserman, A · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Deep video portraits
Kim, H., Garrido, P., Tewari, A., Xu, W., Thies, J., Niessner, M., Pérez, P., Richardt, C., Zollhöfer, M., and Theobalt, C · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A · 2018
Earlier work this paper cites.
Ganimation: Anatomically-aware facial animation from a single image
Pumarola, A., Agudo, A., Martinez, A. M., Sanfeliu, A., and Moreno-Noguer, F · 2018
Earlier work this paper cites.
Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set
Deng, Y., Yang, J., Xu, S., Chen, D., Jia, Y., and Tong, X · 2019
Earlier work this paper cites.
Fvd: A new metric for video generation
Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M · 2020
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2020
Earlier work this paper cites.
A lip sync expert is all you need for speech to lip generation in the wild
Prajwal, K., Mukhopadhyay, R., Namboodiri, V. P., and Jawahar, C · 2020
Earlier work this paper cites.
Blindly assess image quality in the wild guided by a self-adaptive hyper network
Su, S., Yan, Q., Zhu, Y., Zhang, C., Ge, X., Sun, J., and Zhang, Y · 2020
Earlier work this paper cites.
Realistic speech-driven facial animation with gans
Vougioukas, K., Petridis, S., and Pantic, M · 2020
Earlier work this paper cites.
Makelttalk: speaker-aware talking-head animation
Zhou, Y., Han, X., Shechtman, E., Echevarria, J., Kalogerakis, E., and Li, D · 2020
Cited alongside, same era.
Sample and computation redistribution for efficient face detection
Guo, J., Deng, J., Lattas, A., and Zafeiriou, S · 2021
Cited alongside, same era.
Audio-driven emotional video portraits
Ji, X., Zhou, H., Wang, K., Wu, W., Loy, C. C., Cao, X., and Xu, F · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhang, Z., Li, L., Ding, Y., and Fan, C · 2021
Cited alongside, same era.
Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation
Zhang, W., Cun, X., Wang, X., Zhang, Y., Shen, X., Guo, Y., Shan, Y., and Wang, F · 2023
Later among the works it cites.
Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions
Chen, Z., Cao, J., Chen, Z., Li, Y., and Ma, C · 2024
Later among the works it cites.
Liveportrait: Efficient portrait animation with stitching and retargeting control
Guo, J., Zhang, D., Liu, X., Zhong, Z., Zhang, Y., Wan, P., and Zhang, D · 2024
Later among the works it cites.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Hu, L · 2024
Later among the works it cites.
Loopy: Taming audio-driven portrait avatar with long-term motion dependency
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Zhou, H., Sun, Y., Wu, W., Loy, C. C., Wang, X., and Liu, Z · 2021
Cited alongside, same era.
Efficient geometry-aware 3d generative adversarial networks
Chan, E. R., Lin, C. Z., Chan, M. A., Nagano, K., Pan, B., De Mello, S., Gallo, O., Guibas, L. J., Tremblay, J., Khamis, S., et al · 2022
Cited alongside, same era.
Expressive talking head generation with granular audio-visual control
Liang, B., Pan, Y., Guo, Z., Zhou, H., Hong, Z., Han, X., Han, J., Liu, J., Ding, E., and Wang, J · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Celebv-hq: A large-scale video facial attributes dataset
Zhu, H., Wu, W., Zhu, W., Jiang, L., Tang, S., Zhang, L., Liu, Z., and Loy, C. C · 2022
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Liang, Z., Wang, Y., Qiao, Y., Agrawala, M., Lin, D., and Dai, B · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Cited alongside, same era.
Jiang, J., Liang, C., Yang, J., Lin, G., Zhong, T., and Zheng, Y · 2024
Later among the works it cites.
Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation
Ma, Y., Liu, H., Wang, H., Pan, H., He, Y., Yuan, J., Zeng, A., Cai, C., Shum, H.-Y., Liu, W., et al · 2024
Later among the works it cites.
Synctalk: The devil is in the synchronization for talking head synthesis
Peng, Z., Hu, W., Shi, Y., Zhu, X., Zhang, X., Zhao, H., He, J., Liu, H., and Fan, Z · 2024
Later among the works it cites.
Diffused heads: Diffusion models beat gans on talking-face generation
Stypułkowski, M., Vougioukas, K., He, S., Zikeba, M., Petridis, S., and Pantic, M · 2024
Later among the works it cites.
Flowvqtalker: High-quality emotional talking face generation through normalizing flow and quantization
Tan, S., Ji, B., and Pan, Y · 2024
Later among the works it cites.
Tian, L., Wang, Q., Zhang, B., and Bo, L · 2024
Later among the works it cites.
V-express: Conditional dropout for progressive training of portrait video generation
Wang, C., Tian, K., Zhang, J., Guan, Y., Luo, F., Shen, F., Jiang, Z., Gu, Q., Han, X., and Yang, W · 2024
Later among the works it cites.
Aniportrait: Audio-driven synthesis of photorealistic portrait animation
Wei, H., Yang, Z., and Wang, Z · 2024
Later among the works it cites.
Hallo: Hierarchical audio-driven visual synthesis for portrait image animation
Xu, M., Li, H., Su, Q., Shang, H., Zhang, L., Liu, C., Wang, J., Van Gool, L., Yao, Y., and Zhu, S · 2024
Later among the works it cites.
Yang, S., Li, H., Wu, J., Jing, M., Li, L., Ji, R., Liang, J., Fan, H., and Wang, J · 2024
Later among the works it cites.
Real3d-portrait: One-shot realistic 3d talking portrait synthesis
Ye, Z., Zhong, T., Ren, Y., Yang, J., Li, W., Huang, J., Jiang, Z., He, J., Huang, R., Liu, J., et al · 2024
Later among the works it cites.
MEMO: Memory-guided and emotion-aware talking video generation, 2024
Zheng, L., Zhang, Y., Guo, H. A., Pan, J., Tan, Z., Lu, J., Tang, C., An, B., and YAN, S · 2024
Later among the works it cites.