Fetching the paper…
Reading the bibliography…
Text-guided video prediction (TVP) involves predicting the motion of future frames from the initial frame according to an instruction, which has wide applications in virtual reality, robotics, and content creation.
Estimation of non-normalized statistical models by score matching
A. Hyvärinen and P. Dayan · 2005
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
F. Chollet · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. Van Den Oord, O. Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2018
Earlier work this paper cites.
Stochastic adversarial video prediction
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine · 2018
Earlier work this paper cites.
Mocogan: Decomposing motion and content for video generation
S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Earlier work this paper cites.
Adversarial video generation on complex datasets
A. Clark, J. Donahue, and K. Simonyan · 2019
Earlier work this paper cites.
Lower dimensional kernels for video discriminators
E. Kahembwe and S. Ramamoorthy · 2020
Earlier work this paper cites.
Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal gan
M. Saito, S. Saito, M. Koyama, and S. Kobayashi · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Earlier work this paper cites.
Vivit: A video vision transformer
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid · 2021
Earlier work this paper cites.
Is space-time attention all you need for video understanding?
G. Bertasius, H. Wang, and L. Torresani · 2021
Earlier work this paper cites.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2021
Earlier work this paper cites.
Latent neural differential equations for video generation
C. Gordon and N. Parde · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
A good image generator is what you need for high-resolution video synthesis
Y. Tian, J. Ren, M. Chai, K. Olszewski, X. Peng, D. N. Metaxas, and S. Tulyakov · 2021
Earlier work this paper cites.
Videogpt: Video generation using vq-vae and transformers
W. Yan, Y. Zhang, P. Abbeel, and A. Srinivas · 2021
Earlier work this paper cites.
Simvp: Simpler yet better video prediction
Z. Gao, C. Tan, L. Wu, and S. Z. Li · 2022
Earlier work this paper cites.
Long video generation with time-agnostic vqgan and time-sensitive transformer
S. Ge, T. Hayes, H. Yang, X. Yin, G. Pang, D. Jacobs, J.-B. Huang, and D. Parikh · 2022
Earlier work this paper cites.
Latent video diffusion models for high-fidelity video generation with arbitrary lengths
Y. He, T. Yang, Y. Zhang, Y. Shan, and Q. Chen · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, et al · 2022
Earlier work this paper cites.
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang · 2022
Cited alongside, same era.
Make it move: controllable image-to-video generation with text descriptions
Y. Hu, C. Luo, and Z. Chen · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Cited alongside, same era.
Video swin transformer
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu · 2022
Cited alongside, same era.
Sdedit: Image synthesis and editing with stochastic differential equations
C. Meng, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon · 2022
Cited alongside, same era.
On aliased resizing and surprising subtleties in gan evaluation
Flowzero: Zero-shot text-to-video synthesis with llm-driven dynamic scene syntax
Y. Lu, L. Zhu, H. Fan, and Y. Yang · 2023
Later among the works it cites.
Videofusion: Decomposed diffusion models for high-quality video generation
Z. Luo, D. Chen, Y. Zhang, Y. Huang, L. Wang, Y. Shen, D. Zhao, J. Zhou, and T. Tan · 2023
Later among the works it cites.
Gpt4motion: Scripting physical motions in text-to-video generation via blender-oriented gpt planning
J. Lv, Y. Huang, M. Yan, J. Huang, J. Liu, Y. Liu, Y. Wen, X. Chen, and S. Chen · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation
L. Qu, S. Wu, H. Fei, L. Nie, and T.-S. Chua · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Parmar, R. Zhang, and J.-Y. Zhu · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Cited alongside, same era.
Mcvd-masked conditional video diffusion for prediction, generation, and interpolation
V. Voleti, A. Jolicoeur-Martineau, and C. Pal · 2022
Cited alongside, same era.
Generating videos with dynamics-aware implicit generative adversarial networks
S. Yu, J. Tack, S. Mo, H. Kim, J. Kim, J.-W. Ha, and J. Shin · 2022
Cited alongside, same era.
Magicvideo: Efficient video generation with latent diffusion models
D. Zhou, W. Wang, H. Yan, W. Lv, Y. Zhu, and J. Feng · 2022
Cited alongside, same era.
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al · 2023
Later among the works it cites.
Open-vclip: Transforming clip to an open-vocabulary video model via interpolated weight optimization
Z. Weng, X. Yang, A. Li, Z. Wu, and Y.-G. Jiang · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
J. Z. Wu, Y. Ge, X. Wang, W. Lei, Y. Gu, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou · 2023
Later among the works it cites.
Next-gpt: Any-to-any multimodal llm
S. Wu, H. Fei, L. Qu, W. Ji, and T.-S. Chua · 2023
Later among the works it cites.
Svformer: Semi-supervised video transformer for action recognition
Z. Xing, Q. Dai, H. Hu, J. Chen, Z. Wu, and Y.-G. Jiang · 2023
Later among the works it cites.
Vidiff: Translating videos via multi-modal instructions with diffusion models
Z. Xing, Q. Dai, Z. Zhang, H. Zhang, H. Hu, Z. Wu, and Y.-G. Jiang · 2023
Later among the works it cites.
A survey on video diffusion models
Z. Xing, Q. Feng, H. Chen, Q. Dai, H. Hu, H. Xu, Z. Wu, and Y.-G. Jiang · 2023
Later among the works it cites.
Magvit: Masked generative video transformer
L. Yu, Y. Cheng, K. Sohn, J. Lezama, H. Zhang, H. Chang, A. G. Hauptmann, M.-H. Yang, Y. Hao, I. Essa, et al · 2023
Later among the works it cites.
Video probabilistic diffusion models in projected latent space
S. Yu, K. Sohn, S. Kim, and J. Shin · 2023
Later among the works it cites.
Make pixels dance: High-dynamic video generation
Y. Zeng, G. Wei, J. Zheng, J. Zou, Y. Wei, Y. Zhang, and H. Li · 2023
Later among the works it cites.
Adadiff: Adaptive step selection for fast diffusion
H. Zhang, Z. Wu, Z. Xing, J. Shao, and Y.-G. Jiang · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
L. Zhang and M. Agrawala · 2023
Later among the works it cites.
Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024
H. Chen, Y. Zhang, X. Cun, M. Xia, X. Wang, C. Weng, and Y. Shan · 2024
Closest in time.
Layoutgpt: Compositional visual planning and generation with large language models
W. Feng, W. Zhu, T.-j. Fu, V. Jampani, A. Akula, X. He, S. Basu, X. E. Wang, and W. Y. Wang · 2024
Closest in time.
Seer: Language instructed video prediction with latent diffusion models
X. Gu, C. Wen, W. Ye, J. Song, and Y. Gao · 2024
Closest in time.
Free-bloom: Zero-shot text-to-video generator with llm director and ldm animator
H. Huang, Y. Feng, C. Shi, L. Xu, J. Yu, and S. Yang · 2024
Closest in time.
Vstar: Generative temporal nursing for longer dynamic video synthesis
Y. Li, W. Beluch, M. Keuper, D. Zhang, and A. Khoreva · 2024
Closest in time.
Sora, 2024
OpenAI · 2024
Closest in time.
Omnivid: A generative framework for universal video understanding
J. Wang, D. Chen, C. Luo, B. He, L. Yuan, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Microcinema: A divide-and-conquer approach for text-to-video generation
Y. Wang, J. Bao, W. Weng, R. Feng, D. Yin, T. Yang, J. Zhang, Q. Dai, Z. Zhao, C. Wang, K. Qiu, Y. Yuan, C. Tang, X. Sun, C. Luo, and B. Guo · 2024
Closest in time.
Simda: Simple diffusion adapter for efficient video generation
Z. Xing, Q. Dai, H. Hu, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Poseanimate: Zero-shot high fidelity pose controllable character animation
B. Zhu, F. Wang, T. Lu, P. Liu, J. Su, J. Liu, Y. Zhang, Z. Wu, Y.-G. Jiang, and G.-J. Qi · 2024
Closest in time.