Fetching the paper…
Reading the bibliography…
The recent wave of AI-generated content (AIGC) has witnessed substantial success in computer vision, with the diffusion model playing a crucial role in this achievement.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Mean squared error: Love it or leave it? a new look at signal fidelity measures
Z. Wang and A. C. Bovik · 2009
Earlier work this paper cites.
Learning to recognize objects in egocentric activities
A. Fathi, X. Ren, and J. M. Rehg · 2011
Earlier work this paper cites.
A kernel two-sample test
A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
S. Stein and S. J. McKenna · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
H. Kuehne, A. Arslan, and T. Serre · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell · 2017
Earlier work this paper cites.
Realtime multi-person 2d pose estimation using part affinity fields
Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
Carla: An open urban driving simulator
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun · 2017
Earlier work this paper cites.
Self-supervised visual planning with temporal skip connections
F. Ebert, C. Finn, A. X. Lee, and S. Levine · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Earlier work this paper cites.
The thumos challenge on action recognition for videos “in the wild”
H. Idrees, A. R. Zamir, Y.-G. Jiang, A. Gorban, I. Laptev, R. Sukthankar, and M. Shah · 2017
Earlier work this paper cites.
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles · 2017
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
J. Pont-Tuset, F. Perazzi, S. Caelles, P. Arbeláez, A. Sorkine-Hornung, and L. Van Gool · 2017
Earlier work this paper cites.
Movie description
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. Van Den Oord, O. Vinyals, et al · 2017
Earlier work this paper cites.
A short note about kinetics-600
J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2018
Earlier work this paper cites.
Ebsynth: Fast example-based image synthesis and style transfer, 2018
O. Jamriska · 2018
Earlier work this paper cites.
Future frame prediction for anomaly detection–a new baseline
W. Liu, W. Luo, D. Lian, and S. Gao · 2018
Earlier work this paper cites.
How2: a large-scale dataset for multimodal language understanding
R. Sanabria, O. Caglayan, S. Palaskar, D. Elliott, L. Barrault, L. Specia, and F. Metze · 2018
Earlier work this paper cites.
Real-world anomaly detection in surveillance videos
W. Sultani, C. Chen, and M. Shah · 2018
Earlier work this paper cites.
Mocogan: Decomposing motion and content for video generation
S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Earlier work this paper cites.
Video-to-video synthesis
T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, G. Liu, A. Tao, J. Kautz, and B. Catanzaro · 2018
Earlier work this paper cites.
Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks
W. Xiong, W. Luo, L. Ma, W. Liu, and J. Luo · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
L. Zhou, C. Xu, and J. Corso · 2018
Earlier work this paper cites.
An analysis of evaluation metrics of gans
H. Alqahtani, M. Kavakli-Thorne, G. Kumar, and F. SBSSTC · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
T. Karras, S. Laine, and T. Aila · 2019
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding
J. Lin, C. Gan, and S. Han · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic · 2019
Earlier work this paper cites.
First order motion model for image animation
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Y. Song and S. Ermon · 2019
Earlier work this paper cites.
Fvd: A new metric for video generation
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2019
Earlier work this paper cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Earlier work this paper cites.
Generative adversarial networks
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Analyzing and improving the image quality of stylegan
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila · 2020
Earlier work this paper cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun · 2020
Earlier work this paper cites.
Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal gan
M. Saito, S. Saito, M. Koyama, and S. Kobayashi · 2020
Earlier work this paper cites.
Improved techniques for training score-based generative models
Y. Song and S. Ermon · 2020
Earlier work this paper cites.
Learning video representations from textual web supervision
J. C. Stroud, Z. Lu, C. Sun, J. Deng, R. Sukthankar, C. Schmid, and D. A. Ross · 2020
Earlier work this paper cites.
Deep high-resolution representation learning for visual recognition
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y. Zhao, D. Liu, Y. Mu, M. Tan, X. Wang, et al · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
M. Bain, A. Nagrani, G. Varol, and A. Zisserman · 2021
Earlier work this paper cites.
Conditional image generation with score-based diffusion models
G. Batzolis, J. Stanczuk, C.-B. Schönlieb, and C. Etmann · 2021
Earlier work this paper cites.
Ilvr: Conditioning method for denoising diffusion probabilistic models
J. Choi, S. Kim, Y. Jeong, Y. Gwon, and S. Yoon · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2021
Earlier work this paper cites.
Learning high fidelity depths of dressed humans by watching social media dance videos
Y. Jafarian and H. S. Park · 2021
Earlier work this paper cites.
Alias-free generative adversarial networks
T. Karras, M. Aittala, S. Laine, E. Härkönen, J. Hellsten, J. Lehtinen, and T. Aila · 2021
Earlier work this paper cites.
Ccvs: context-aware controllable video synthesis
G. Le Moing, J. Ponce, and C. Schmid · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Earlier work this paper cites.
Diffusion probabilistic models for 3d point cloud generation
S. Luo and W. Hu · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Earlier work this paper cites.
Maximum likelihood training of score-based diffusion models
Y. Song, C. Durkan, I. Murray, and S. Ermon · 2021
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2021
Earlier work this paper cites.
A good image generator is what you need for high-resolution video synthesis
Y. Tian, J. Ren, M. Chai, K. Olszewski, X. Peng, D. N. Metaxas, and S. Tulyakov · 2021
Earlier work this paper cites.
Score-based generative modeling in latent space
A. Vahdat, K. Kreis, and J. Kautz · 2021
Earlier work this paper cites.
Videogpt: Video generation using vq-vae and transformers
W. Yan, Y. Zhang, P. Abbeel, and A. Srinivas · 2021
Earlier work this paper cites.
Merlot: Multimodal neural script knowledge models
R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi · 2021
Earlier work this paper cites.
Videolt: Large-scale long-tailed video recognition
X. Zhang, Z. Wu, Z. Weng, H. Fu, J. Chen, Y.-G. Jiang, and L. S. Davis · 2021
Earlier work this paper cites.
Blended diffusion for text-driven editing of natural images
O. Avrahami, D. Lischinski, and O. Fried · 2022
Earlier work this paper cites.
ediffi: Text-to-image diffusion models with an ensemble of expert denoisers
Y. Balaji, S. Nah, X. Huang, A. Vahdat, J. Song, K. Kreis, M. Aittala, T. Aila, S. Laine, B. Catanzaro, et al · 2022
Earlier work this paper cites.
Text2live: Text-driven layered image and video editing
O. Bar-Tal, D. Ofri-Amar, R. Fridman, Y. Kasten, and T. Dekel · 2022
Earlier work this paper cites.
Unleashing transformers: Parallel token prediction with discrete absorbing diffusion for fast high-resolution image generation from vector-quantized codes
S. Bond-Taylor, P. Hessey, H. Sasaki, T. P. Breckon, and C. G. Willcocks · 2022
Earlier work this paper cites.
Generating long videos of dynamic scenes
T. Brooks, J. Hellsten, M. Aittala, T.-C. Wang, T. Aila, J. Lehtinen, M.-Y. Liu, A. Efros, and T. Karras · 2022
Earlier work this paper cites.
Diffusiondet: Diffusion model for object detection
S. Chen, P. Sun, Y. Song, and P. Luo · 2022
Earlier work this paper cites.
A generalist framework for panoptic segmentation of images and videos
T. Chen, L. Li, S. Saxena, G. Hinton, and D. J. Fleet · 2022
Earlier work this paper cites.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
M. Ding, W. Zheng, W. Hong, and J. Tang · 2022
Earlier work this paper cites.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2022
Earlier work this paper cites.
Vision-language pre-training: Basics, recent advances, and future trends
Z. Gan, L. Li, C. Li, L. Wang, Z. Liu, J. Gao, et al · 2022
Earlier work this paper cites.
Long video generation with time-agnostic vqgan and time-sensitive transformer
S. Ge, T. Hayes, H. Yang, X. Yin, G. Pang, D. Jacobs, J.-B. Huang, and D. Parikh · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Earlier work this paper cites.
Vector quantized diffusion model for text-to-image synthesis
S. Gu, D. Chen, J. Bao, F. Wen, B. Zhang, D. Chen, L. Yuan, and B. Guo · 2022
Earlier work this paper cites.
Diffusioninst: Diffusion model for instance segmentation
Z. Gu, H. Chen, Z. Xu, J. Lan, C. Meng, and W. Wang · 2022
Earlier work this paper cites.
Flexible diffusion modeling of long videos
W. Harvey, S. Naderiparizi, V. Masrani, C. Weilbach, and F. Wood · 2022
Earlier work this paper cites.
Latent video diffusion models for high-fidelity video generation with arbitrary lengths
Y. He, T. Yang, Y. Zhang, Y. Shan, and Q. Chen · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, et al · 2022
Earlier work this paper cites.
Cascaded diffusion models for high fidelity image generation
J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans · 2022
Earlier work this paper cites.
Video diffusion models
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet · 2022
Earlier work this paper cites.
Diffusion models for video prediction and infilling
T. Höppe, A. Mehrjou, S. Bauer, D. Nielsen, and A. Dittadi · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Earlier work this paper cites.
Denoising diffusion restoration models
B. Kawar, M. Elad, S. Ermon, and J. Song · 2022
Earlier work this paper cites.
Sdedit: Image synthesis and editing with stochastic differential equations
C. Meng, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon · 2022
Earlier work this paper cites.
Midjourney., 2022
Midjourney · 2022
Earlier work this paper cites.
Null-text inversion for editing real images using guided diffusion models
R. Mokady, A. Hertz, K. Aberman, Y. Pritch, and D. Cohen-Or · 2022
Earlier work this paper cites.
Learning audio-video modalities from image captions
A. Nagrani, P. H. Seo, B. Seybold, A. Hauth, S. Manen, C. Sun, and C. Schmid · 2022
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2022
Cited alongside, same era.
Sinfusion: Training diffusion models on a single image or video
Y. Nikankin, N. Haim, and M. Irani · 2022
Cited alongside, same era.
Chatgpt: A large-scale generative model for conversational ai, 2022
OpenAI · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
On aliased resizing and surprising subtleties in gan evaluation
G. Parmar, R. Zhang, and J.-Y. Zhu · 2022
Cited alongside, same era.
Follow your pose: Pose-guided text-to-video generation using pose-free videos
Y. Ma, Y. He, X. Cun, X. Wang, Y. Shan, X. Li, and Q. Chen · 2023
Closest in time.
Vidm: Video implicit diffusion models
K. Mei and V. Patel · 2023
Closest in time.
On distillation of guided diffusion models
C. Meng, R. Rombach, R. Gao, D. Kingma, S. Ermon, J. Ho, and T. Salimans · 2023
Closest in time.
Dreamix: Video diffusion models are general video editors
E. Molad, E. Horwitz, D. Valevski, A. R. Acha, Y. Matias, Y. Pritch, Y. Leviathan, and Y. Hoshen · 2023
Closest in time.
Difftad: Temporal action detection with proposal denoising diffusion
S. Nag, X. Zhu, J. Deng, Y.-Z. Song, and T. Xiang · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Cited alongside, same era.
Image super-resolution via iterative refinement
C. Saharia, J. Ho, W. Chan, T. Salimans, D. J. Fleet, and M. Norouzi · 2022
Cited alongside, same era.
Progressive distillation for fast sampling of diffusion models
T. Salimans and J. Ho · 2022
Cited alongside, same era.
J. Nam, G. Lee, S. Kim, H. Kim, H. Cho, S. Kim, and S. Kim · 2023
Closest in time.
Conditional image-to-video generation with latent flow diffusion models
H. Ni, C. Shi, K. Li, S. X. Huang, and M. R. Min · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Closest in time.
Fatezero: Fusing attentions for zero-shot text-based video editing
C. Qi, X. Cun, Y. Zhang, C. Lei, X. Wang, Y. Shan, and Q. Chen · 2023
Closest in time.
Instructvid2vid: Controllable video editing with natural language instructions
B. Qin, J. Li, S. Tang, T.-S. Chua, and Y. Zhuang · 2023
Closest in time.
Dancing avatar: Pose and text-guided human motion videos synthesis with image diffusion model
B. Qin, W. Ye, Q. Yu, S. Tang, and Y. Zhuang · 2023
Closest in time.
Feature-conditioned cascaded video diffusion models for precise echocardiogram synthesis
H. Reynaud, M. Qiao, M. Dombrowski, T. Day, R. Razavi, A. Gomez, P. Leeson, and B. Kainz · 2023
Closest in time.
Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation
L. Ruan, Y. Ma, H. Yang, H. He, B. Liu, J. Fu, N. J. Yuan, Q. Jin, and B. Guo · 2023
Closest in time.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman · 2023
Closest in time.
Edit-a-video: Single video editing with object-aware consistency
C. Shin, H. Kim, C. H. Lee, S.-g. Lee, and S. Yoon · 2023
Closest in time.
Make-a-video: Text-to-video generation without text-video data
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al · 2023
Closest in time.
Any-to-any generation via composable diffusion
Z. Tang, Z. Yang, C. Zhu, M. Zeng, and M. Bansal · 2023
Closest in time.
Plug-and-play diffusion features for text-driven image-to-image translation
N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel · 2023
Closest in time.
Exploring diffusion models for unsupervised video anomaly detection
A. O. Tur, N. Dall’Asen, C. Beyan, and E. Ricci · 2023
Closest in time.
Unsupervised video anomaly detection with diffusion models conditioned on compact motion representations
A. O. Tur, N. Dall’Asen, C. Beyan, and E. Ricci · 2023
Closest in time.
Gen-l-video: Multi-text to long video generation via temporal co-denoising
F.-Y. Wang, W. Chen, G. Song, H.-J. Ye, Y. Liu, and H. Li · 2023
Closest in time.
Dformer: Diffusion-guided transformer for universal image segmentation
H. Wang, J. Cao, R. M. Anwer, J. Xie, F. S. Khan, and Y. Pang · 2023
Closest in time.
Pdpp: Projected diffusion for procedure planning in instructional videos
H. Wang, Y. Wu, S. Guo, and L. Wang · 2023
Closest in time.
Modelscope text-to-video technical report
J. Wang, H. Yuan, D. Chen, Y. Zhang, X. Wang, and S. Zhang · 2023
Closest in time.
Disco: Disentangled control for referring human dance generation in real world
T. Wang, L. Li, K. Lin, C.-C. Lin, Z. Yang, H. Zhang, Z. Liu, and L. Wang · 2023
Closest in time.
Zero-shot video editing using off-the-shelf image diffusion models
W. Wang, K. Xie, Z. Liu, H. Chen, Y. Cao, X. Wang, and C. Shen · 2023
Closest in time.
Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation
W. Wang, H. Yang, Z. Tuo, H. He, J. Zhu, J. Fu, and J. Liu · 2023
Closest in time.
Videocomposer: Compositional video synthesis with motion controllability
X. Wang, H. Yuan, S. Zhang, D. Chen, J. Wang, Y. Zhang, Y. Shen, D. Zhao, and J. Zhou · 2023
Closest in time.
Videolcm: Video latent consistency model
X. Wang, S. Zhang, H. Zhang, Y. Liu, Y. Zhang, C. Gao, and N. Sang · 2023
Closest in time.
Lavie: High-quality video generation with cascaded latent diffusion models
Y. Wang, X. Chen, X. Ma, S. Zhou, Z. Huang, Y. Wang, C. Yang, Y. He, J. Yu, P. Yang, et al · 2023
Closest in time.
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Y. Wang, Y. He, Y. Li, K. Li, J. Yu, X. Ma, X. Chen, Y. Wang, P. Luo, Z. Liu, et al · 2023
Closest in time.
Styleinv: A temporal style modulated inversion network for unconditional video generation
Y. Wang, L. Jiang, and C. C. Loy · 2023
Closest in time.
Edit temporal-consistent videos with image diffusion model
Y. Wang, Y. Li, X. Liu, A. Dai, A. Chan, and Z. Cui · 2023
Closest in time.
Leo: Generative latent image animator for human video synthesis
Y. Wang, X. Ma, X. Chen, A. Dantcheva, B. Dai, and Y. Qiao · 2023
Closest in time.
Motionctrl: A unified and flexible motion controller for video generation
Z. Wang, Z. Yuan, X. Wang, T. Chen, M. Xia, P. Luo, and Y. Shan · 2023
Closest in time.
Ai-generated content (aigc): A survey
J. Wu, W. Gan, Z. Chen, S. Wan, and H. Lin · 2023
Closest in time.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
J. Z. Wu, Y. Ge, X. Wang, W. Lei, Y. Gu, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou · 2023
Closest in time.
Next-gpt: Any-to-any multimodal llm
S. Wu, H. Fei, L. Qu, W. Ji, and T.-S. Chua · 2023
Closest in time.
Make-your-video: Customized video generation using textual and structural guidance
J. Xing, M. Xia, Y. Liu, Y. Zhang, Y. Zhang, Y. He, H. Liu, H. Chen, X. Cun, X. Wang, et al · 2023
Closest in time.
Dynamicrafter: Animating open-domain images with video diffusion priors
J. Xing, M. Xia, Y. Zhang, H. Chen, X. Wang, T.-T. Wong, and Y. Shan · 2023
Closest in time.
Vidiff: Translating videos via multi-modal instructions with diffusion models
Z. Xing, Q. Dai, Z. Zhang, H. Zhang, H. Hu, Z. Wu, and Y.-G. Jiang · 2023
Closest in time.
Multimodal-driven talking face generation via a unified diffusion-based generator
C. Xu, S. Zhu, J. Zhu, T. Huang, J. Zhang, Y. Tai, and Y. Liu · 2023
Closest in time.
Magicprop: Diffusion-based video editing via motion-aware appearance propagation
H. Yan, J. H. Liew, L. Mai, S. Lin, and J. Feng · 2023
Closest in time.
Probabilistic adaptation of text-to-video models
M. Yang, Y. Du, B. Dai, D. Schuurmans, J. B. Tenenbaum, and P. Abbeel · 2023
Closest in time.
Video diffusion models with local-global context guidance
S. Yang, L. Zhang, Y. Liu, Z. Jiang, and Y. He · 2023
Closest in time.
Rerender a video: Zero-shot text-guided video-to-video translation
S. Yang, Y. Zhou, Z. Liu, and C. C. Loy · 2023
Closest in time.
Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory
S. Yin, C. Wu, J. Liang, J. Shi, H. Li, G. Ming, and N. Duan · 2023
Closest in time.
Nuwa-xl: Diffusion over diffusion for extremely long video generation
S. Yin, C. Wu, H. Yang, J. Wang, X. Wang, M. Ni, Z. Yang, L. Li, S. Liu, F. Yang, et al · 2023
Closest in time.
Long-term rhythmic video soundtracker
J. Yu, Y. Wang, X. Chen, X. Sun, and Y. Qiao · 2023
Closest in time.
Celebv-text: A large-scale facial text-video dataset
J. Yu, H. Zhu, L. Jiang, C. C. Loy, W. Cai, and W. Wu · 2023
Closest in time.
Video probabilistic diffusion models in projected latent space
S. Yu, K. Sohn, S. Kim, and J. Shin · 2023
Closest in time.
Make pixels dance: High-dynamic video generation
Y. Zeng, G. Wei, J. Zheng, J. Zou, Y. Wei, Y. Zhang, and H. Li · 2023
Closest in time.
Multimodal image synthesis and editing: A survey and taxonomy
F. Zhan, Y. Yu, R. Wu, J. Zhang, S. Lu, L. Liu, A. Kortylewski, C. Theobalt, and E. Xing · 2023
Closest in time.
Text-to-image diffusion model in generative ai: A survey
C. Zhang, C. Zhang, M. Zhang, and I. S. Kweon · 2023
Closest in time.
A complete survey on generative ai (aigc): Is chatgpt from gpt-4 to gpt-5 all you need?
C. Zhang, C. Zhang, S. Zheng, Y. Qiao, C. Li, M. Zhang, S. K. Dam, C. M. Thwal, Y. L. Tun, L. L. Huy, et al · 2023
Closest in time.
Show-1: Marrying pixel and latent diffusion models for text-to-video generation
D. J. Zhang, J. Z. Wu, J.-W. Liu, R. Zhao, L. Ran, Y. Gu, D. Gao, and M. Z. Shou · 2023
Closest in time.
Adadiff: Adaptive step selection for fast diffusion
H. Zhang, Z. Wu, Z. Xing, J. Shao, and Y.-G. Jiang · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
L. Zhang and M. Agrawala · 2023
Closest in time.
Controlvideo: Training-free controllable text-to-video generation
Y. Zhang, Y. Wei, D. Jiang, X. Zhang, W. Zuo, and Q. Tian · 2023
Closest in time.
Towards consistent video editing with text-to-image diffusion models
Z. Zhang, B. Li, X. Nie, C. Han, T. Guo, and L. Liu · 2023
Closest in time.
Diffusionvmr: Diffusion model for video moment retrieval
H. Zhao, K. Q. Lin, R. Yan, and Z. Li · 2023
Closest in time.
Controlvideo: Adding conditional control for one shot text-to-video editing
M. Zhao, R. Wang, F. Bao, C. Li, and J. Zhu · 2023
Closest in time.
Make-a-protagonist: Generic video editing with an ensemble of experts
Y. Zhao, E. Xie, L. Hong, Z. Li, and G. H. Lee · 2023
Closest in time.
Refined semantic enhancement towards frequency diffusion for video captioning
X. Zhong, Z. Li, S. Chen, K. Jiang, C. Chen, and M. Ye · 2023
Closest in time.
Vision+ language applications: A survey
Y. Zhou and N. Shimada · 2023
Closest in time.
J. Zhu, H. Yang, H. He, W. Wang, Z. Tuo, W.-H. Cheng, L. Gao, J. Song, and J. Fu · 2023
Closest in time.
Gentron: Diffusion transformers for image and video generation
S. Chen, M. Xu, J. Ren, Y. Cong, S. He, Y. Xie, A. Sinha, P. Luo, T. Xiang, and J.-M. Perez-Rua · 2024
Closest in time.
Panda-70m: Captioning 70m videos with multiple cross-modality teachers
T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H.-w. Chao, B. E. Jeon, Y. Fang, H.-Y. Lee, J. Ren, M.-H. Yang, et al · 2024
Closest in time.
Diffutoon: High-resolution editable toon shading via diffusion models
Z. Duan, C. Wang, C. Chen, W. Qian, and J. Huang · 2024
Closest in time.
Aigcbench: Comprehensive evaluation of image-to-video content generated by ai
F. Fan, C. Luo, J. Zhan, and W. Gao · 2024
Closest in time.
Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing, 2024
H. Fei, S. Wu, H. Zhang, T.-S. Chua, and S. Yan · 2024
Closest in time.
Fdgaussian: Fast gaussian splatting from single image via geometric-aware diffusion model
Q. Feng, Z. Xing, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Matten: Video generation with mamba-attention
Y. Gao, J. Huang, X. Sun, Z. Jie, Y. Zhong, and L. Ma · 2024
Closest in time.
On the content bias in fréchet video distance
S. Ge, A. Mahapatra, G. Parmar, J.-Y. Zhu, and J.-B. Huang · 2024
Closest in time.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
L. Hu · 2024
Closest in time.
Vbench: Comprehensive benchmark suite for video generative models
Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, et al · 2024
Closest in time.
Stream: Spatio-temporal evaluation and analysis metric for video generative models
P. J. Kim, S. Kim, and J. Yoo · 2024
Closest in time.
Grid diffusion models for text-to-video generation
T. Lee, S. Kwon, and T. Kim · 2024
Closest in time.
Flowvid: Taming imperfect optical flows for consistent video-to-video synthesis
F. Liang, B. Wu, J. Wang, L. Yu, K. Li, Y. Zhao, I. Misra, J.-B. Huang, P. Zhang, P. Vajda, et al · 2024
Closest in time.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.
Evalcrafter: Benchmarking and evaluating large video generation models
Y. Liu, X. Cun, X. Liu, X. Wang, Y. Zhang, H. Chen, Y. Liu, T. Zeng, R. Chan, and Y. Shan · 2024
Closest in time.
Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation
Y. Liu, L. Li, S. Ren, R. Gao, S. Li, S. Chen, X. Sun, and L. Hou · 2024
Closest in time.
Latte: Latent diffusion transformer for video generation
X. Ma, Y. Wang, G. Jia, X. Chen, Z. Liu, Y.-F. Li, C. Chen, and Y. Qiao · 2024
Closest in time.
Snap video: Scaled spatiotemporal transformers for text-to-video synthesis
W. Menapace, A. Siarohin, I. Skorokhodov, E. Deyneka, T.-S. Chen, A. Kag, Y. Fang, A. Stoliar, E. Ricci, J. Ren, et al · 2024
Closest in time.
Scaling diffusion mamba with bidirectional ssms for efficient image and video generation
S. Mo and Y. Tian · 2024
Closest in time.
Ti2v-zero: Zero-shot image conditioning for text-to-video diffusion models
H. Ni, B. Egger, S. Lohit, A. Cherian, Y. Wang, T. Koike-Akino, S. X. Huang, and T. K. Marks · 2024
Closest in time.
Sora, 2024
OpenAI · 2024
Closest in time.
Flexifilm: Long video generation with flexible conditions
Y. Ouyang, H. Zhao, G. Wang, et al · 2024
Closest in time.
L. Tian, Q. Wang, B. Zhang, and L. Bo · 2024
Closest in time.
Motioneditor: Editing video motion via content-aware diffusion
S. Tu, Q. Dai, Z.-Q. Cheng, H. Hu, X. Han, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Omnitokenizer: A joint image-video tokenizer for visual generation
J. Wang, Y. Jiang, Z. Yuan, B. Peng, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Microcinema: A divide-and-conquer approach for text-to-video generation
Y. Wang, J. Bao, W. Weng, R. Feng, D. Yin, T. Yang, J. Zhang, Q. Dai, Z. Zhao, C. Wang, et al · 2024
Closest in time.
Dreamvideo: Composing your dream videos with customized subject and motion
Y. Wei, S. Zhang, Z. Qing, H. Yuan, Z. Liu, Y. Liu, Y. Zhang, J. Zhou, and H. Shan · 2024
Closest in time.
Genrec: Unifying video generation and recognition with diffusion models
Z. Weng, X. Yang, Z. Xing, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Fairy: Fast parallelized instruction-guided video-to-video synthesis
B. Wu, C.-Y. Chuang, X. Wang, Y. Jia, K. Krishnakumar, T. Xiao, F. Liang, L. Yu, and P. Vajda · 2024
Closest in time.
Towards a better metric for text-to-video generation
J. Z. Wu, G. Fang, H. Wu, X. Wang, Y. Ge, X. Cun, D. J. Zhang, J.-W. Liu, Y. Gu, R. Zhao, et al · 2024
Closest in time.
Simda: Simple diffusion adapter for efficient video generation
Z. Xing, Q. Dai, H. Hu, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Aid: Adapting image2video diffusion models for instruction-guided video prediction
Z. Xing, Q. Dai, Z. Weng, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Magicanimate: Temporally consistent human image animation using diffusion model
Z. Xu, J. Zhang, J. H. Liew, H. Yan, J.-W. Liu, C. Zhang, J. Feng, and M. Z. Shou · 2024
Closest in time.
Efficient video diffusion models via content-frame motion-latent decomposition
S. Yu, W. Nie, D.-A. Huang, B. Li, J. Shin, and A. Anandkumar · 2024
Closest in time.
Avid: Any-length video inpainting with diffusion model
Z. Zhang, B. Wu, X. Wang, Y. Luo, L. Zhang, Y. Zhao, P. Vajda, D. Metaxas, and L. Yu · 2024
Closest in time.
Poseanimate: Zero-shot high fidelity pose controllable character animation
B. Zhu, F. Wang, T. Lu, P. Liu, J. Su, J. Liu, Y. Zhang, Z. Wu, Y.-G. Jiang, and G.-J. Qi · 2024
Closest in time.