Fetching the paper…
Reading the bibliography…
Text-to-video (T2V) generation has been recently enabled by transformer-based diffusion models, but current T2V models lack capabilities in adhering to the real-world common knowledge and physical rules, due to their limited understanding of physical realism and deficiency in temporal modeling.
Blender, an open source design tool: Advances and integration in the architectural production pipeline
Theodoros Dounas and Alexandros Sigalas · 2009
Earlier work this paper cites.
Unrealcv: Connecting computer vision to unreal engine
Weichao Qiu and Alan Yuille · 2016
Earlier work this paper cites.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, et al · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang · 2022
Earlier work this paper cites.
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al · 2022
Earlier work this paper cites.
Towards understanding chain-of-thought prompting: An empirical study of what matters
Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, and Huan Sun · 2022
Earlier work this paper cites.
A systematic survey of prompt engineering on vision-language foundation models, 2023
Jindong Gu, Zhen Han, Shuo Chen, Ahmad Beirami, Bailan He, Gengyuan Zhang, Ruotong Liao, Yao Qin, Volker Tresp, and Philip Torr · 2023
Earlier work this paper cites.
Pika labs, 2023
Mellis Inc · 2023
Earlier work this paper cites.
Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning
Han Lin, Abhay Zala, Jaemin Cho, and Mohit Bansal · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Earlier work this paper cites.
Learning interactive real-world simulators
Mengjiao Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Dale Schuurmans, and Pieter Abbeel · 2023
Cited alongside, same era.
Ai video generation expert discusses the technology’s rapid advances—and its current limitations, 2024
Luke Auburn · 2024
Cited alongside, same era.
Videophy: Evaluating physical commonsense for video generation
Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang, and Aditya Grover · 2024
Cited alongside, same era.
Video generation models as world simulators, 2024
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al · 2024
Cited alongside, same era.
Videocrafter2: Overcoming data limitations for high-quality video diffusion models
Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan · 2024
Unity real-time development platform, 2024
Unity Technologies · 2024
Closest in time.
A.i. can now create lifelike videos. can you tell what’s real?, 2024
Stuart A. Thompson · 2024
Closest in time.
Self-correcting llm-controlled diffusion models
Tsung-Han Wu, Long Lian, Joseph E Gonzalez, Boyi Li, and Trevor Darrell · 2024
Closest in time.
Wentao Zhang, Junliang Guo, Tianyu He, Li Zhao, Linli Xu, and Jiang Bian · 2024
Closest in time.
Open-sora: Democratizing efficient video production for all, 2024
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, et al · 2024
Cited alongside, same era.
Upbge: an open-source, 3d game engine forked from the old blender game engine, 2024
Blender Foundation · 2024
Cited alongside, same era.
Game engines for immersive visualization: Using unreal engine beyond entertainment
Marcel Krüger, David Gilbert, Torsten W. Kuhlen, and Tim Gerrits · 2024
Cited alongside, same era.
Towards world simulator: Crafting physical commonsense-based benchmark for video generation
Fanqing Meng, Jiaqi Liao, Xinyu Tan, Wenqi Shao, Quanfeng Lu, Kaipeng Zhang, Yu Cheng, Dianqi Li, Yu Qiao, and Ping Luo · 2024
Cited alongside, same era.
A systematic survey of prompt engineering in large language models: Techniques and applications, 2024
Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha, Vinija Jain, Samrat Mondal, and Aman Chadha · 2024
Cited alongside, same era.
Freezeasguard: Mitigating illegal adaptation of diffusion models via selective tensor freezing
Kai Huang, Haoming Wang, and Wei Gao
Cited in the paper.
Towards green ai in fine-tuning large language models via adaptive backpropagation
Kai Huang, Hanyun Yin, Heng Huang, and Wei Gao
Cited in the paper.
Hanxin Zhu, Tianyu He, Anni Tang, Junliang Guo, Zhibo Chen, and Jiang Bian · 2024
Closest in time.
Photorealistic video generation with diffusion models
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Fei-Fei Li, Irfan Essa, Lu Jiang, and José Lezama · 2025
Closest in time.
Modality plug-and-play: Runtime modality adaptation in LLM-driven autonomous mobile systems
Kai Huang, Xiangyu Yin, Heng Huang, and Wei Gao · 2025
Closest in time.
Physgen: Rigid-body physics-grounded image-to-video generation
Shaowei Liu, Zhongzheng Ren, Saurabh Gupta, and Shenlong Wang · 2025
Closest in time.