Fetching the paper…
Reading the bibliography…
Character animation is a transformative field in computer graphics and vision, enabling dynamic and realistic video animations from static images.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., and Black, M. J · 2015
Earlier work this paper cites.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Earlier work this paper cites.
Realtime multi-person 2d pose estimation using part affinity fields
Cao, Z., Simon, T., Wei, S.-E., and Sheikh, Y · 2017
Earlier work this paper cites.
Densepose: Dense human pose estimation in the wild
Güler, R. A., Neverova, N., and Kokkinos, I · 2018
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2021
Earlier work this paper cites.
ediffi: Text-to-image diffusion models with an ensemble of expert denoisers
Balaji, Y., Nah, S., Huang, X., Vahdat, A., Song, J., Kreis, K., Aittala, M., Aila, T., Laine, S., Catanzaro, B., et al · 2022
Earlier work this paper cites.
Video diffusion models
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Cited alongside, same era.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2022
Cited alongside, same era.
Structure and content-guided video synthesis with diffusion models
Esser, P., Chiu, J., Atighehchian, P., Granskog, J., and Germanidis, A · 2023
Fatezero: Fusing attentions for zero-shot text-based video editing
QI, C., Cun, X., Zhang, Y., Lei, C., Wang, X., Shan, Y., and Chen, Q · 2023
Later among the works it cites.
Advancing pose-guided image synthesis with progressive conditional diffusion models
Shen, F., Ye, H., Zhang, J., Wang, C., Han, X., and Wei, Y · 2023
Later among the works it cites.
Objectstitch: Object compositing with diffusion model
Song, Y., Zhang, Z., Lin, Z., Cohen, S., Price, B., Zhang, J., Kim, S. Y., and Aliaga, D · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z · 2023
Later among the works it cites.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Ye, H., Zhang, J., Liu, S., Han, X., and Yang, W · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., and Dai, B · 2023
Cited alongside, same era.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Hu, L., Gao, X., Zhang, P., Sun, K., Zhang, B., and Bo, L · 2023
Cited alongside, same era.
Composer: creative and controllable image synthesis with composable conditions
Huang, L., Chen, D., Liu, Y., Shen, Y., Zhao, D., and Zhou, J · 2023
Cited alongside, same era.
Dreampose: Fashion video synthesis with stable diffusion
Karras, J., Holynski, A., Wang, T.-C., and Kemelmacher-Shlizerman, I · 2023
Cited alongside, same era.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
Khachatryan, L., Movsisyan, A., Tadevosyan, V., Henschel, R., Wang, Z., Navasardyan, S., and Shi, H · 2023
Cited alongside, same era.
Mou, C., Wang, X., Xie, L., Zhang, J., Qi, Z., Shan, Y., and Qie, X · 2023
Cited alongside, same era.
G2l: Semantically aligned and uniform video grounding via geodesic and game theory
Li, H., Cao, M., Cheng, X., Li, Y., Zhu, Z., and Zou, Y
Cited in the paper.
Adding conditional control to text-to-image diffusion models
Zhang, L., Rao, A., and Agrawala, M · 2023
Later among the works it cites.
Tryondiffusion: A tale of two unets
Zhu, L., Yang, D., Zhu, T., Reda, F., Chan, W., Saharia, C., Norouzi, M., and Kemelmacher-Shlizerman, I · 2023
Later among the works it cites.
Exploiting auxiliary caption for video grounding
Li, H., Cao, M., Cheng, X., Li, Y., Zhu, Z., and Zou, Y · 2024
Closest in time.
Coarse-to-fine latent diffusion for pose-guided person image synthesis
Lu, Y., Zhang, M., Ma, A. J., Xie, X., and Lai, J · 2024
Closest in time.
Enhancing fine-grained multi-modal alignment via adapters: A parameter-efficient training framework for referring image segmentation
Xu, Z., Huang, J., Liu, T., Liu, Y., Han, H., Yuan, K., and Li, X · 2024
Closest in time.
Yang, S., Xu, Z., Xue, H., Cheng, Y., Huang, S., Gong, M., and Wu, Z · 2024
Closest in time.