Fetching the paper…
Reading the bibliography…
Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics.
MoCoGAN: Decomposing motion and content for video generation
Tulyakov, S., Liu, M.-Y., Yang, X., and Kautz, J · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow
Teed, Z. and Deng, J · 2020
Earlier work this paper cites.
Diffusion models beat GANs on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Earlier work this paper cites.
Compositional visual generation with composable diffusion models
Liu, N., Li, S., Du, Y., Torralba, A., and Tenenbaum, J. B · 2022
Earlier work this paper cites.
SDEdit: Guided image synthesis and editing with stochastic differential equations
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
UL2: Unifying language learning paradigms
Tay, Y., Dehghani, M., Tran, V. Q., Garcia, X., Wei, J., Wang, X., Chung, H. W., Shakeri, S., Bahri, D., Schuster, T., et al · 2022
Earlier work this paper cites.
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Xue, L., Barua, A., Constant, N., Al-Rfou, R., Narang, S., Kale, M., Roberts, A., and Raffel, C · 2022
Earlier work this paper cites.
Latent-Shift: Latent diffusion with temporal shift for efficient text-to-video generation
An, J., Zhang, S., Yang, H., Gupta, S., Huang, J.-B., Luo, J., and Yin, X · 2023
Earlier work this paper cites.
InstructPix2Pix: Learning to follow image editing instructions
Brooks, T., Holynski, A., and Efros, A. A · 2023
Earlier work this paper cites.
Flatten: optical flow-guided attention for consistent text-to-video editing
Cong, Y., Xu, M., Simon, C., Chen, S., Ren, J., Xie, Y., Perez-Rua, J.-M., Rosenhahn, B., Xiang, T., and He, S · 2023
Earlier work this paper cites.
Emu: Enhancing image generation models using photogenic needles in a haystack
Dai, X., Hou, J., Ma, C.-Y., Tsai, S., Wang, J., Wang, R., Zhang, P., Vandenhende, S., Wang, X., Dubey, A., et al · 2023
Cited alongside, same era.
AnimateDiff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., and Dai, B · 2023
Cited alongside, same era.
Flow matching for generative modeling
Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M · 2023
Cited alongside, same era.
Hyperhuman: Hyper-realistic human generation with latent structural diffusion
Liu, X., Ren, J., Siarohin, A., Skorokhodov, I., Li, Y., Lin, D., Liu, X., Liu, Z., and Tulyakov, S · 2023
Cited alongside, same era.
Trailblazer: Trajectory control for diffusion-based video generation, 2023
Motion prompting: Controlling video generation with motion trajectories, 2024
Geng, D., Herrmann, C., Hur, J., Cole, F., Zhang, S., Pfaff, T., Lopez-Guevara, T., Doersch, C., Aytar, Y., Rubinstein, M., Sun, C., Wang, O., Owens, A., and Sun, D · 2024
Later among the works it cites.
Emu video: Factorizing text-to-video generation by explicit image conditioning
Girdhar, R., Singh, M., Brown, A., Duval, Q., Azadi, S., Rambhatla, S. S., Shah, A., Yin, X., Parikh, D., and Misra, I · 2024
Later among the works it cites.
Ltx-video: Realtime video latent diffusion
HaCohen, Y., Chiprut, N., Brazowski, B., Shalem, D., Moshe, D., Richardson, E., Levin, E., Shiran, G., Zabari, N., Gordon, O., Panet, P., Weissbuch, S., Kulikov, V., Bitterman, Y., Melumian, Z., and Bibi, O · 2024
Later among the works it cites.
VBench: Comprehensive benchmark suite for video generative models
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., Wang, Y., Chen, X., Wang, L., Lin, D., Qiao, Y., and Liu, Z · 2024
Later among the works it cites.
How far is video generation from world model: A physical law perspective, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ma, W.-D. K., Lewis, J. P., and Kleijn, W. B · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Hierarchical spatio-temporal decoupling for text-to-video generation, 2023
Qing, Z., Zhang, S., Wang, J., Wang, X., Wei, Y., Zhang, Y., Gao, C., and Sang, N · 2023
Cited alongside, same era.
DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K · 2023
Cited alongside, same era.
Make-A-Video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., Parikh, D., Gupta, S., and Taigman, Y · 2023
Cited alongside, same era.
ModelScope text-to-video technical report
Wang, J., Yuan, H., Chen, D., Zhang, Y., Wang, X., and Zhang, S · 2023
Cited alongside, same era.
Tune-A-Video: One-shot tuning of image diffusion models for text-to-video generation
Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z · 2023
Cited alongside, same era.
Xu, H., Xie, S., Tan, X. E., Huang, P.-Y., Howes, R., Sharma, V., Li, S.-W., Ghosh, G., Zettlemoyer, L., and Feichtenhofer, C · 2023
Cited alongside, same era.
Kang, B., Yue, Y., Lu, R., Lin, Z., Zhao, Y., Wang, K., Huang, G., and Feng, J · 2024
Later among the works it cites.
Kling AI, 2024
KlingAI · 2024
Later among the works it cites.
Motioncraft: Physics-based zero-shot video generation
Montanaro, A., Savant Aira, L., Aiello, E., Valsesia, D., and Magli, E · 2024
Later among the works it cites.
Dall-E 3, 2024
OpenAI · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models, 2024
Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.-Y., Chuang, C.-Y., Yan, D., Choudhary, D., Wang, D., Sethi, G., Pang, G., Ma, H., Misra, I., Hou, J., Wang, J., Jagadeesh, K., Li, K., Zhang, L., Singh, M., Williamson, M., Le, M., Yu, M., Singh, M. K., Zhang, P., Vajda, P., Duval, Q., Girdhar, R., Sumbaly, R., Rambhatla, S. S., Tsai, S., Azadi, S., Datta, S., Chen, S., Bell, S., Ramaswamy, S., Sheynin, S., Bhattacharya, S., Motwani, S., Xu, T., Li, T., Hou, T., Hsu, W.-N., Yin, X., Dai, X., Taigman, Y., Luo, Y., Liu, Y.-C., Wu, Y.-C., Zhao, Y., Kirstain, Y., He, Z., He, Z., Pumarola, A., Thabet, A., Sanakoyeu, A., Mallya, A., Guo, B., Araya, B., Kerr, B., Wood, C., Liu, C., Peng, C., Vengertsev, D., Schonfeld, E., Blanchard, E., Juefei-Xu, F., Nord, F., Liang, J., Hoffman, J., Kohler, J., Fire, K., Sivakumar, K., Chen, L., Yu, L., Gao, L., Georgopoulos, M., Moritz, R., Sampson, S. K., Li, S., Parmeggiani, S., Fine, S., Fowler, T., Petrovic, V., and Du, Y · 2024
Later among the works it cites.
Enhancing motion in text-to-video generation with decomposed encoding and conditioning, 2024
Ruan, P., Wang, P., Saxena, D., Cao, J., and Shi, Y · 2024
Later among the works it cites.
Gen-3 Alpha, 2024
RunwayML · 2024
Later among the works it cites.
Decouple content and motion for conditional image-to-video generation
Shen, C., Gan, Y., Chen, C., Zhu, X., Cheng, L., Gao, T., and Wang, J · 2024
Later among the works it cites.
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024
Shi, X., Huang, Z., Wang, F.-Y., Bian, W., Li, D., Zhang, Y., Zhang, M., Cheung, K. C., See, S., Qin, H., Dai, J., and Li, H · 2024
Later among the works it cites.
Motif: Making text count in image animation with motion focal loss, 2024
Wang, S., Azadi, S., Girdhar, R., Rambhatla, S., Sun, C., and Yin, X · 2024
Later among the works it cites.