Fetching the paper…
Reading the bibliography…
Text-to-video (T2V) synthesis has gained increasing attention in the community, in which the recently emerged diffusion models (DMs) have promisingly shown stronger performance than the past approaches.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir D. Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian J. Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Video pixel networks
Nal Kalchbrenner, Aäron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Encoding sentences with graph convolutional networks for semantic role labeling
Diego Marcheggiani and Ivan Titov · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Towards high resolution video generation with progressive growing of sliced wasserstein gans
Dinesh Acharya, Zhiwu Huang, Danda Pani Paudel, and Luc Van Gool · 2018
Earlier work this paper cites.
Image generation from scene graphs
Justin Johnson, Agrim Gupta, and Li Fei-Fei · 2018
Earlier work this paper cites.
Video generation from text
Yitong Li, Martin Min, Dinghan Shen, David Carlson, and Lawrence Carin · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphaël Marinier, Marcin Michalski, and Sylvain Gelly · 2018
Earlier work this paper cites.
Graph attention networks
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio · 2018
Earlier work this paper cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He · 2018
Earlier work this paper cites.
Adversarial video generation on complex datasets
Aidan Clark, Jeff Donahue, and Karen Simonyan · 2019
Earlier work this paper cites.
Recurrent space-time graph neural networks
Andrei Liviu Nicolicioiu, Iulia Duta, and Marius Leordeanu · 2019
Earlier work this paper cites.
Markov decision process for video generation
Vladyslav Yushchenko, Nikita Araslanov, and Stefan Roth · 2019
Earlier work this paper cites.
DM-GAN: dynamic memory generative adversarial networks for text-to-image synthesis
Minfeng Zhu, Pingbo Pan, Wei Chen, and Yi Yang · 2019
Earlier work this paper cites.
Latent neural differential equations for video generation
Cade Gordon and Natalie Parde · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2020
Earlier work this paper cites.
Lower dimensional kernels for video discriminators
Emmanuel Kahembwe and Subramanian Ramamoorthy · 2020
Earlier work this paper cites.
Videoflow: A conditional flow-based model for stochastic video generation
Manoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn, Sergey Levine, Laurent Dinh, and Durk Kingma · 2020
Earlier work this paper cites.
Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal GAN
Masaki Saito, Shunta Saito, Masanori Koyama, and Sosuke Kobayashi · 2020
Earlier work this paper cites.
Unbiased scene graph generation from biased training
Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang · 2020
Earlier work this paper cites.
DF-GAN: deep fusion generative adversarial networks for text-to-image synthesis
Ming Tao, Hao Tang, Songsong Wu, Nicu Sebe, Fei Wu, and Xiao-Yuan Jing · 2020
Earlier work this paper cites.
Scaling autoregressive video models
Dirk Weissenborn, Oscar Täckström, and Jakob Uszkoreit · 2020
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Cited alongside, same era.
A flow-based latent state generative model of neural population responses to natural images
Mohammad Bashiri, Edgar Y. Walker, Konstantin-Klemens Lurz, Akshay Jagadish, Taliah Muhammad, Zhiwei Ding, Zhuokun Ding, Andreas S. Tolias, and Fabian H. Sinz · 2021
Cited alongside, same era.
Stylevideogan: A temporal generative model using a pretrained stylegan
Gereon Fox, Ayush Tewari, Mohamed Elgharib, and Christian Theobalt · 2021
Cited alongside, same era.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Cited alongside, same era.
Temporal shift GAN for large scale video generation
Generating videos with dynamics-aware implicit generative adversarial networks
Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin · 2022
Later among the works it cites.
Magicvideo: Efficient video generation with latent diffusion models
Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng · 2022
Later among the works it cites.
Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation
Jie An, Songyang Zhang, Harry Yang, Sonal Gupta, Jia-Bin Huang, Jiebo Luo, and Xi Yin · 2023
Closest in time.
Align your latents: High-resolution video synthesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis · 2023
Closest in time.
Structure and content-guided video synthesis with diffusion models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrés Muñoz, Mohammadreza Zolfaghari, Max Argus, and Thomas Brox · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Cited alongside, same era.
A good image generator is what you need for high-resolution video synthesis
Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N. Metaxas, and Sergey Tulyakov · 2021
Cited alongside, same era.
Videogpt: Video generation using VQ-VAE and transformers
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas · 2021
Cited alongside, same era.
Image-to-image retrieval by learning similarity between scene graphs
Sangwoong Yoon, Woo-Young Kang, Sungwook Jeon, SeongEun Lee, Changjin Han, Jonghun Park, and Eun-Sol Kim · 2021
Cited alongside, same era.
Multi-scale 2d temporal adjacency networks for moment localization with natural language
Songyang Zhang, Houwen Peng, Jianlong Fu, Yijuan Lu, and Jiebo Luo · 2021
Cited alongside, same era.
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis · 2023
Closest in time.
Preserve your own correlation: A noise prior for video diffusion models
Songwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon, Andrew Tao, Bryan Catanzaro, David Jacobs, Jia-Bin Huang, Ming-Yu Liu, and Yogesh Balaji · 2023
Closest in time.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection
Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu · 2023
Closest in time.
Can chatgpt boost artistic creation: The need of imaginative intelligence for parallel art
Chao Guo, Yue Lu, Yong Dou, and Fei-Yue Wang · 2023
Closest in time.
Videogen: A reference-guided latent diffusion approach for high definition text-to-video generation
Xin Li, Wenqing Chu, Ye Wu, Weihang Yuan, Fanglong Liu, Qi Zhang, Fu Li, Haocheng Feng, Errui Ding, and Jingdong Wang · 2023
Closest in time.
Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning
Han Lin, Abhay Zala, Jaemin Cho, and Mohit Bansal · 2023
Closest in time.
Ed-t2v: An efficient training framework for diffusion-based text-to-video generation
Jiawei Liu, Weining Wang, Wei Liu, Qian He, and Jing Liu · 2023
Closest in time.
Decomposed diffusion models for high-quality video generation
Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jinren Zhou, and Tieniu Tan · 2023
Closest in time.
Videofusion: Decomposed diffusion models for high-quality video generation
Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan · 2023
Closest in time.
Conditional image-to-video generation with latent flow diffusion models
Haomiao Ni, Changhao Shi, Kai Li, Sharon X. Huang, and Martin Renqiang Min · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation
Leigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie, and Tat-Seng Chua · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Chatgpt empowered long-step robot control in various environments: A case application
Naoki Wake, Atsushi Kanehira, Kazuhiro Sasabuchi, Jun Takamatsu, and Katsushi Ikeuchi · 2023
Closest in time.
Modelscope text-to-video technical report
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang · 2023
Closest in time.
Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation
Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu · 2023
Closest in time.
Videocomposer: Compositional video synthesis with motion controllability
Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou · 2023
Closest in time.
Lavie: High-quality video generation with cascaded latent diffusion models
Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al · 2023
Closest in time.
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinyuan Chen, Yaohui Wang, Ping Luo, Ziwei Liu, et al · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan · 2023
Closest in time.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou · 2023
Closest in time.
Simda: Simple diffusion adapter for efficient video generation
Zhen Xing, Qi Dai, Han Hu, Zuxuan Wu, and Yu-Gang Jiang · 2023
Closest in time.
Video probabilistic diffusion models in projected latent space
Sihyun Yu, Kihyuk Sohn, Subin Kim, and Jinwoo Shin · 2023
Closest in time.
Show-1: Marrying pixel and latent diffusion models for text-to-video generation
David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao, Lingmin Ran, Yuchao Gu, Difei Gao, and Mike Zheng Shou · 2023
Closest in time.
Motiondirector: Motion customization of text-to-video diffusion models
Rui Zhao, Yuchao Gu, Jay Zhangjie Wu, David Junhao Zhang, Jiawei Liu, Weijia Wu, Jussi Keppo, and Mike Zheng Shou · 2023
Closest in time.
Imagine that! abstract-to-intricate text-to-image synthesis with scene graph hallucination diffusion
Shengqiong Wu, Hao Fei, Hanwang Zhang, and Tat-Seng Chua · 2024
Closest in time.