Fetching the paper…
Reading the bibliography…
Recent large-scale pre-trained diffusion models have demonstrated a powerful generative ability to produce high-quality videos from detailed text descriptions.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2017 · 2017
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Bain, M.; Nagrani, A.; Varol, G.; and Zisserman, A. 2021 · 2021
Earlier work this paper cites.
Leveraging intra-domain knowledge to strengthen cross-domain crowd counting
Cai, Y.; Chen, L.; Ma, Z.; Lu, C.; Wang, C.; and He, G. 2021 · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J.; Holtzman, A.; Forbes, M.; Bras, R. L.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
Global Representation Guided Adaptive Fusion Network for Stable Video Crowd Counting
Cai, Y.; Ma, Z.; Lu, C.; Wang, C.; and He, G. 2022 · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
Ho, J.; Chan, W.; Saharia, C.; Whang, J.; Gao, R.; Gritsenko, A.; Kingma, D. P.; Poole, B.; Norouzi, M.; Fleet, D. J.; et al. 2022 · 2022
Earlier work this paper cites.
Simple Open-Vocabulary Object Detection with Vision Transformers
Minderer, M.; Gritsenko, A.; Stone, A.; Neumann, M.; Weissenborn, D.; Dosovitskiy, A.; Mahendran, A.; Arnab, A.; Dehghani, M.; Shen, Z.; Wang, X.; Zhai, X.; Kipf, T.; and Houlsby, N. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Earlier work this paper cites.
Advancing high-resolution video-language representation with large-scale video transcriptions
Xue, H.; Hang, T.; Zeng, Y.; Sun, Y.; Liu, B.; Yang, H.; Fu, J.; and Guo, B. 2022 · 2022
Earlier work this paper cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y.; English, Z.; Voleti, V.; Letts, A.; et al. 2023 · 2023
Earlier work this paper cites.
Explicit invariant feature induced cross-domain crowd counting
Cai, Y.; Chen, L.; Guan, H.; Lin, S.; Lu, C.; Wang, C.; and He, G. 2023 · 2023
Cited alongside, same era.
Diffusion self-guidance for controllable image generation
Epstein, D.; Jabri, A.; Poole, B.; Efros, A. A.; and Holynski, A. 2023 · 2023
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y.; Yang, C.; Rao, A.; Wang, Y.; Qiao, Y.; Lin, D.; and Dai, B. 2023 · 2023
Cited alongside, same era.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Kirstain, Y.; Polyak, A.; Singer, U.; Matiana, S.; Penna, J.; and Levy, O. 2023 · 2023
Cited alongside, same era.
TrailBlazer: Trajectory Control for Diffusion-Based Video Generation
Ma, W.-D. K.; Lewis, J. P.; and Kleijn, W. 2023 · 2023
Show-1: Marrying pixel and latent diffusion models for text-to-video generation
Zhang, D. J.; Wu, J. Z.; Liu, J.-W.; Zhao, R.; Ran, L.; Gu, Y.; Gao, D.; and Shou, M. Z. 2023 · 2023
Later among the works it cites.
Adding Conditional Control to Text-to-Image Diffusion Models
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Later among the works it cites.
Multi-Prototype Space Learning for Commonsense-Based Scene Graph Generation
Chen, L.; Song, Y.; Cai, Y.; Lu, J.; Li, Y.; Xie, Y.; Wang, C.; and He, G. 2024 · 2024
Closest in time.
Peekaboo: Interactive video generation via masked-diffusion
Jain, Y.; Nasery, A.; Vineet, V.; and Behl, H. 2024 · 2024
Closest in time.
FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models
Qiu, H.; Chen, Z.; Wang, Z.; He, Y.; Xia, M.; and Liu, Z. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
GVQA: Learning to Answer Questions about Graphs with Visualizations via Knowledge Base
Song, S.; Chen, J.; Li, C.; and Wang, C. 2023a · 2023
Cited alongside, same era.
ZeroScope
Sterling, S. 2023 · 2023
Cited alongside, same era.
Multi-Source Templates Learning for Real-Time Aerial Tracking
Sun, Y.; Li, Y.; and Wang, C. 2023 · 2023
Cited alongside, same era.
Contrastive Pseudo Learning for Open-World DeepFake Attribution
Sun, Z.; Chen, S.; Yao, T.; Yin, B.; Yi, R.; Ding, S.; and Ma, L. 2023 · 2023
Cited alongside, same era.
CVPR 2023 Text Guided Video Editing Competition
Wu, J. Z.; Li, X.; Gao, D.; Dong, Z.; Bai, J.; Singh, A.; Xiang, X.; Li, Y.; Huang, Z.; Sun, Y.; He, R.; Hu, F.; Hu, J.; Huang, H.; Zhu, H.; Cheng, X.; Tang, J.; Shou, M. Z.; Keutzer, K.; and Iandola, F. 2023b · 2023
Cited alongside, same era.
Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion
Xie, J.; Li, Y.; Huang, Y.; Liu, H.; Zhang, W.; Zheng, Y.; and Shou, M. Z. 2023 · 2023
Cited alongside, same era.
Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory
Yin, S.; Wu, C.; Liang, J.; Shi, J.; Li, H.; Ming, G.; and Duan, N. 2023 · 2023
Cited alongside, same era.
Closest in time.
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling
Shi, X.; Huang, Z.; Wang, F.-Y.; Bian, W.; Li, D.; Zhang, Y.; Zhang, M.; Cheung, K. C.; See, S.; Qin, H.; et al. 2024 · 2024
Closest in time.
GraphDecoder: Recovering Diverse Network Graphs From Visualization Images via Attention-Aware Learning
Song, S.; Li, C.; Li, D.; Chen, J.; and Wang, C. 2024 · 2024
Closest in time.
Motionbooth: Motion-aware customized text-to-video generation
Wu, J.; Li, X.; Zeng, Y.; Zhang, J.; Zhou, Q.; Li, Y.; Tong, Y.; and Chen, K. 2024 · 2024
Closest in time.
ClothPPO: A Proximal Policy Optimization Enhancing Framework for Robotic Cloth Manipulation with Observation-Aligned Action Spaces
Yang, L.; Li, Y.; and Chen, L. 2024 · 2024
Closest in time.
Zero-shot controllable image-to-video animation via motion decomposition
Yu, S.; Fang, J. Z.; Zheng, J.; Sigurdsson, G.; Ordonez, V.; Piramuthu, R.; and Bansal, M. 2024 · 2024
Closest in time.
Non-uniform Timestep Sampling: Towards Faster Diffusion Model Training
Zheng, T.; Geng, C.; Jiang, P.; Wan, B.; Zhang, H.; Chen, J.; Wang, J.; and Li, B. 2024a · 2024
Closest in time.
Beta-Tuned Timestep Diffusion Model
Zheng, T.; Jiang, P.; Wan, B.; Zhang, H.; Chen, J.; Wang, J.; and Li, B. 2024b · 2024
Closest in time.