Fetching the paper…
Reading the bibliography…
Recent advances in diffusion models have led to significant progress in audio-driven lip synchronization.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Chung, J. S. and Zisserman, A · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I · 2017
Earlier work this paper cites.
Ava-speech: A densely labeled dataset of speech activity in movies
Chaudhuri, S., Roth, J., Ellis, D. P., Gallagher, A., Kaver, L., Marvin, R., Pantofaru, C., Reale, N., Reid, L. G., Wilson, K., et al · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
A lip sync expert is all you need for speech to lip generation in the wild
Prajwal, K., Mukhopadhyay, R., Namboodiri, V. P., and Jawahar, C · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2020
Earlier work this paper cites.
Blindly assess image quality in the wild guided by a self-adaptive hyper network
Su, S., Yan, Q., Zhu, Y., Zhang, C., Ge, X., Sun, J., and Zhang, Y · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Earlier work this paper cites.
Pixel difference networks for efficient edge detection
Su, Z., Liu, W., Yu, Z., Hu, D., Liao, Q., Tian, Q., Pietikäinen, M., and Liu, L · 2021
Earlier work this paper cites.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhang, Z., Li, L., Ding, Y., and Fan, C · 2021
Earlier work this paper cites.
Videoretalking: Audio-based lip synchronization for talking head video editing in the wild
Cheng, K., Cun, X., Zhang, Y., Xia, M., Yin, F., Zhu, M., Wang, X., Wang, J., and Wang, N · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Earlier work this paper cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
Ranftl, R., Lasinger, K., Hafner, D., Schindler, K., and Koltun, V · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
The MPEG immersive video standard - current status and future outlook
Vadakital, V. K. M., Dziembowski, A., Lafruit, G., Thudor, F., Lee, G., and Rondao-Alface, P · 2022
Cited alongside, same era.
Vfhq: A high-quality dataset and benchmark for video face super-resolution
Wang, L. X., Zhang, H., Dong, C., Shan, Y., et al · 2022
Cited alongside, same era.
Instructpix2pix: Learning to follow image editing instructions
Brooks, T., Holynski, A., and Efros, A. A · 2023
Cited alongside, same era.
Video generation models as world simulators, 2024
Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A · 2024
Later among the works it cites.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Hu, L · 2024
Later among the works it cites.
Loopy: Taming audio-driven portrait avatar with long-term motion dependency
Jiang, J., Liang, C., Yang, J., Lin, G., Zhong, T., and Zheng, Y · 2024
Later among the works it cites.
Latentsync: Audio conditioned latent diffusion models for lip sync
Li, C., Zhang, C., Xu, W., Xie, J., Feng, W., Peng, B., and Xing, W · 2024
Later among the works it cites.
Latte: Latent diffusion transformer for video generation
Ma, X., Wang, Y., Jia, G., Chen, X., Liu, Z., Li, Y., Chen, C., and Qiao, Y · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Magicdance: Realistic human dance video generation with motions & facial expressions transfer
Chang, D., Shi, Y., Gao, Q., Fu, J., Xu, H., Song, G., Yan, Q., Yang, X., and Soleymani, M · 2023
Cited alongside, same era.
Structure and content-guided video synthesis with diffusion models
Esser, P., Chiu, J., Atighehchian, P., Granskog, J., and Germanidis, A · 2023
Cited alongside, same era.
Emu video: Factorizing text-to-video generation by explicit image conditioning
Girdhar, R., Singh, M., Brown, A., Duval, Q., Azadi, S., Rambhatla, S. S., Shah, A., Yin, X., Parikh, D., and Misra, I · 2023
Cited alongside, same era.
Stylesync: High-fidelity generalized and personalized lip sync in style-based generator
Guan, J., Zhang, Z., Zhou, H., Hu, T., Wang, K., He, D., Feng, H., Liu, J., Ding, E., Liu, Z., et al · 2023
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., and Dai, B · 2023
Cited alongside, same era.
Ground-a-video: Zero-shot grounded video editing using text-to-image diffusion models
Jeong, H. and Ye, J. C · 2023
Cited alongside, same era.
Magicedit: High-fidelity and temporally coherent video editing
Liew, J. H., Yan, H., Zhang, J., Xu, Z., and Feng, J · 2023
Cited alongside, same era.
Later among the works it cites.
Diff2lip: Audio conditioned diffusion models for lip-synchronization
Mukhopadhyay, S., Suri, S., Gadde, R. T., and Shrivastava, A · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models
Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.-Y., Chuang, C.-Y., et al · 2024
Later among the works it cites.
Lane segmentation refinement with diffusion models
Ruiz, A., Melnik, A., Wang, D., and Ritter, H · 2024
Later among the works it cites.
Audio-driven high-resolution seamless talking head video editing via stylegan
Su, J., Liu, K., Chen, L., Yao, J., Liu, Q., and Lv, D · 2024
Later among the works it cites.
Multitalk: Enhancing 3d talking head generation across languages with multilingual video dataset
Sung-Bin, K., Chae-Yeon, L., Son, G., Hyun-Bin, O., Ju, J., Nam, S., and Oh, T.-H · 2024
Later among the works it cites.
Style2talker: High-resolution talking head generation with emotion style and art style
Tan, S., Ji, B., and Pan, Y · 2024
Later among the works it cites.
Ladtalk: Latent denoising for synthesizing talking head videos with high frequency details
Yang, J., Wang, X., Wang, W., Li, G., Fang, Q., Yuan, R., Wang, T., and Fan, J. Z · 2024
Later among the works it cites.
Musetalk: Real-time high quality lip synchronization with latent space inpainting
Zhang, Y., Liu, M., Chen, Z., Wu, B., Zeng, Y., Zhan, C., He, Y., Huang, J., and Zhou, W · 2024
Later among the works it cites.
Open-sora: Democratizing efficient video production for all, Mar 2024
Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., and You, Y · 2024
Later among the works it cites.
High-fidelity and lip-synced talking face synthesis via landmark-based diffusion model
Zhong, W., Lin, J., Chen, P., Lin, L., and Li, G · 2024
Later among the works it cites.
Infp: Audio-driven interactive head generation in dyadic conversations
Zhu, Y., Zhang, L., Rong, Z., Hu, T., Liang, S., and Ge, Z · 2024
Later among the works it cites.
Edtalk: Efficient disentanglement for emotional talking head synthesis
Tan, S., Ji, B., Bi, M., and Pan, Y · 2025
Closest in time.
Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions
Tian, L., Wang, Q., Zhang, B., and Bo, L · 2025
Closest in time.