Fetching the paper…
Reading the bibliography…
We introduce PhysMotion, a novel framework that leverages principled physics-based simulations to guide intermediate 3D representations generated from a single image and input conditions (e.g., applied force and torque), producing high-quality, physically plausible video generation.
Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation. In Computer Vision and Pattern Recognition (CVPR) . 1921–1930
Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. 2023 · 1930
Earlier work this paper cites.
The OpenCV Library
G. Bradski. 2000 · 2000
Earlier work this paper cites.
A material point method for snow simulation
Alexey Stomakhin, Craig Schroeder, Lawrence Chai, Joseph Teran, and Andrew Selle. 2013 · 2013
Earlier work this paper cites.
Optimization integrator for large time steps
Theodore F Gast, Craig Schroeder, Alexey Stomakhin, Chenfanfu Jiang, and Joseph M Teran. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning . PMLR, 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
The material point method for simulating continuum materials
Chenfanfu Jiang, Craig Schroeder, Joseph Teran, Alexey Stomakhin, and Andrew Selle. 2016 · 2016
Earlier work this paper cites.
XPBD: position-based simulation of compliant constrained dynamics. In Proceedings of the 9th International Conference on Motion in Games . 49–54
Miles Macklin, Matthias Müller, and Nuttapong Chentanez. 2016 · 2016
Earlier work this paper cites.
Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . Curran Associates Inc., Red Hook, NY, USA, 6000–6010
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Style transfer by relaxed optimal transport and self-similarity. In Computer Vision and Pattern Recognition (CVPR) . 10051–10060
Nicholas Kolkin, Jason Salavon, and Gregory Shakhnarovich. 2019 · 2019
Earlier work this paper cites.
Swapping autoencoder for deep image manipulation
Taesung Park, Jun-Yan Zhu, Oliver Wang, Jingwan Lu, Eli Shechtman, Alexei Efros, and Richard Zhang. 2020 · 2020
Earlier work this paper cites.
Segdiff: Image segmentation with diffusion probabilistic models
Tomer Amit, Tal Shaharbany, Eliya Nachmani, and Lior Wolf. 2021 · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. 2021 · 2021
Earlier work this paper cites.
Putting nerf on a diet: Semantically consistent few-shot view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 5885–5894
Ajay Jain, Matthew Tancik, and Pieter Abbeel. 2021 · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021 · 2021
Earlier work this paper cites.
pixelnerf: Neural radiance fields from one or few images. In Computer Vision and Pattern Recognition (CVPR) . 4578–4587
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. 2021 · 2021
Earlier work this paper cites.
ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers
Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, et al · 2022
Earlier work this paper cites.
Depth-supervised nerf: Fewer views and faster training for free. In Computer Vision and Pattern Recognition (CVPR) . 12882–12891
Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ramanan. 2022 · 2022
Earlier work this paper cites.
Classifier-Free Diffusion Guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
Energetically consistent inelasticity for optimization time integration
Xuan Li, Minchen Li, and Chenfanfu Jiang. 2022 · 2022
Earlier work this paper cites.
Repaint: Inpainting using denoising diffusion probabilistic models. In Computer Vision and Pattern Recognition (CVPR) . 11461–11471
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022 · 2022
Earlier work this paper cites.
Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In Computer Vision and Pattern Recognition (CVPR) . 5480–5490
Michael Niemeyer, Jonathan T Barron, Ben Mildenhall, Mehdi SM Sajjadi, Andreas Geiger, and Noha Radwan. 2022 · 2022
Earlier work this paper cites.
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022 · 2022
Earlier work this paper cites.
DreamBooth: Fine Tuning Text-to-image Diffusion Models for Subject-Driven Generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Pix2Video: Video Editing using Image Diffusion. In International Conference on Computer Vision (ICCV)
Duygu Ceylan, Chun-Hao Huang, and Niloy J. Mitra. 2023 · 2023
Earlier work this paper cites.
Stablevideo: Text-driven consistency-aware diffusion video editing. In International Conference on Computer Vision (ICCV) . 23040–23050
Wenhao Chai, Xun Guo, Gaoang Wang, and Yan Lu. 2023 · 2023
Earlier work this paper cites.
Control-a-video: Controllable text-to-video generation with diffusion models
Weifeng Chen, Yatai Ji, Jie Wu, Hefeng Wu, Pan Xie, Jiashi Li, Xin Xia, Xuefeng Xiao, and Liang Lin. 2023 · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. 2023 · 2023
Earlier work this paper cites.
Expressive text-to-image generation with rich text. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 7545–7556
Songwei Ge, Taesung Park, Jun-Yan Zhu, and Jia-Bin Huang. 2023 · 2023
Earlier work this paper cites.
TokenFlow: Consistent Diffusion Features for Consistent Video Editing
Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. 2023 · 2023
Earlier work this paper cites.
Animate-a-story: Storytelling with retrieval-augmented video generation
Yingqing He, Menghan Xia, Haoxin Chen, Xiaodong Cun, Yuan Gong, Jinbo Xing, Yong Zhang, Xintao Wang, Chao Weng, Ying Shan, et al · 2023
Earlier work this paper cites.
Imagic: Text-based real image editing with diffusion models. In Computer Vision and Pattern Recognition (CVPR) . 6007–6017
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. 2023 · 2023
Earlier work this paper cites.
3D Gaussian Splatting for Real-Time Radiance Field Rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 2023 · 2023
Earlier work this paper cites.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. 2023 · 2023
Cited alongside, same era.
Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model
Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan Sunkavalli, Greg Shakhnarovich, and Sai Bi. 2023b · 2023
Cited alongside, same era.
Xuan Li, Yi-Ling Qiao, Peter Yichen Chen, Krishna Murthy Jatavallabhula, Ming Lin, Chenfanfu Jiang, and Chuang Gan. 2023a · 2023
Cited alongside, same era.
Magicedit: High-fidelity and temporally coherent video editing
Jun Hao Liew, Hanshu Yan, Jianfeng Zhang, Zhongcong Xu, and Jiashi Feng. 2023 · 2023
Cited alongside, same era.
3d diffuser actor: Policy diffusion with 3d scene representations
Tsung-Wei Ke, Nikolaos Gkanatsios, and Katerina Fragkiadaki. 2024 · 2024
Closest in time.
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
Max Ku, Cong Wei, Weiming Ren, Harry Yang, and Wenhu Chen. 2024 · 2024
Closest in time.
DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization
Jiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng, Xin Ning, Jun Zhou, and Lin Gu. 2024d · 2024
Closest in time.
Animate Your Motion: Turning Still Images into Dynamic Videos
Mingxiao Li, Bo Wan, Marie-Francine Moens, and Tinne Tuytelaars. 2024c · 2024
Closest in time.
Phy124: Fast Physics-Driven 4D Content Generation from a Single Image
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Magic3d: High-resolution text-to-3d content creation. In Computer Vision and Pattern Recognition (CVPR) . 300–309
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. 2023 · 2023
Cited alongside, same era.
Zero-1-to-3: Zero-shot one image to 3d object. In Proceedings of the IEEE/CVF international conference on computer vision . 9298–9309
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. 2023 · 2023
Cited alongside, same era.
On distillation of guided diffusion models. In Computer Vision and Pattern Recognition (CVPR) . 14297–14306
Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans. 2023 · 2023
Cited alongside, same era.
Conditional image-to-video generation with latent flow diffusion models. In Computer Vision and Pattern Recognition (CVPR) . 18444–18455
Haomiao Ni, Changhao Shi, Kai Li, Sharon X Huang, and Martin Renqiang Min. 2023 · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023 · 2023
Cited alongside, same era.
Fatezero: Fusing attentions for zero-shot text-based video editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 15932–15942
Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. 2023 · 2023
Cited alongside, same era.
DreamGaussian4D: Generative 4D Gaussian Splatting
Jiawei Ren, Liang Pan, Jiaxiang Tang, Chi Zhang, Ang Cao, Gang Zeng, and Ziwei Liu. 2023 · 2023
Cited alongside, same era.
MVDream: Multi-view Diffusion for 3D Generation
Yichun Shi, Peng Wang, Jianglong Ye, Long Mai, Kejie Li, and Xiao Yang. 2023 · 2023
Cited alongside, same era.
Jiajing Lin, Zhenzhong Wang, Yongjie Hou, Yuzhou Tang, and Min Jiang. 2024 · 2024
Closest in time.
Video-p2p: Video editing with cross-attention control. In Computer Vision and Pattern Recognition (CVPR) . 8599–8608
Shaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin, and Jiaya Jia. 2024 · 2024
Closest in time.
Wonder3d: Single image to 3d using cross-domain diffusion. In Computer Vision and Pattern Recognition (CVPR) . 9970–9980
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al · 2024
Closest in time.
Snap video: Scaled spatiotemporal transformers for text-to-video synthesis. In Computer Vision and Pattern Recognition (CVPR) . 7038–7048
Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov, Ekaterina Deyneka, Tsai-Shien Chen, Anil Kag, Yuwei Fang, Aleksei Stoliar, Elisa Ricci, Jian Ren, et al · 2024
Closest in time.
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
Fanqing Meng, Jiaqi Liao, Xinyu Tan, Wenqi Shao, Quanfeng Lu, Kaipeng Zhang, Yu Cheng, Dianqi Li, Yu Qiao, and Ping Luo. 2024 · 2024
Closest in time.
CoherentGS: Sparse novel view synthesis with coherent 3D Gaussians
Avinash Paliwal, Wei Ye, Jinhui Xiong, Dmytro Kotovenko, Rakesh Ranjan, Vikas Chandra, and Nima Khademi Kalantari. 2024 · 2024
Closest in time.
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling. In ACM SIGGRAPH 2024 Conference Papers . 1–11
Xiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian, Dasong Li, Yi Zhang, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, et al · 2024
Closest in time.
Splatter Image: Ultra-Fast Single-View 3D Reconstruction. In Computer Vision and Pattern Recognition (CVPR)
Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. 2024 · 2024
Closest in time.
LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation
Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. 2024 · 2024
Closest in time.
Motionctrl: A unified and flexible motion controller for video generation. In ACM SIGGRAPH 2024 Conference Papers . 1–11
Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. 2024 · 2024
Closest in time.
Dreamvideo: Composing your dream videos with customized subject and motion. In Computer Vision and Pattern Recognition (CVPR) . 6537–6549
Yujie Wei, Shiwei Zhang, Zhiwu Qing, Hangjie Yuan, Zhiheng Liu, Yu Liu, Yingya Zhang, Jingren Zhou, and Hongming Shan. 2024 · 2024
Closest in time.
latentsplat: Autoencoding variational gaussians for fast generalizable 3d reconstruction
Christopher Wewer, Kevin Raj, Eddy Ilg, Bernt Schiele, and Jan Eric Lenssen. 2024 · 2024
Closest in time.
AniClipart: Clipart Animation with Text-to-Video Priors
Ronghuan Wu, Wanchao Su, Kede Ma, and Jing Liao. 2024b · 2024
Closest in time.
Direct-a-video: Customized video generation with user-directed camera movement and object motion. In ACM SIGGRAPH 2024 Conference Papers . 1–12
Shiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma, Pengfei Wan, Di Zhang, Xiaodong Chen, and Jing Liao. 2024a · 2024
Closest in time.
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al · 2024
Closest in time.
Space-time diffusion features for zero-shot text-driven motion transfer. In Computer Vision and Pattern Recognition (CVPR) . 8466–8476
Danah Yatim, Rafail Fridman, Omer Bar-Tal, Yoni Kasten, and Tali Dekel. 2024 · 2024
Closest in time.
Image Sculpting: Precise Object Editing with 3D Geometry Control
Jiraphon Yenphraphai, Xichen Pan, Sainan Liu, Daniele Panozzo, and Saining Xie. 2024 · 2024
Closest in time.
WonderWorld: Interactive 3D Scene Generation from a Single Image
Hong-Xing Yu, Haoyi Duan, Charles Herrmann, William T Freeman, and Jiajun Wu. 2024 · 2024
Closest in time.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations. In ICRA 2024 Workshop on 3D Visual Representations for Robot Manipulation
Yanjie Ze, Gu Zhang, Kangning Zhang, Chenyuan Hu, Muhan Wang, and Huazhe Xu. 2024 · 2024
Closest in time.
Mimicmotion: High-quality human motion video generation with confidence-aware pose guidance
Yuang Zhang, Jiaxi Gu, Li-Wen Wang, Han Wang, Junqi Cheng, Yuefeng Zhu, and Fangyuan Zou. 2024a · 2024
Closest in time.
Tora: Trajectory-oriented Diffusion Transformer for Video Generation
Zhenghao Zhang, Junchao Liao, Menghao Li, Long Qin, and Weizhi Wang. 2024b · 2024
Closest in time.
A Convex Formulation of Frictional Contact for the Material Point Method and Rigid Bodies
Zeshun Zong, Chenfanfu Jiang, and Xuchen Han. 2024 · 2024
Closest in time.
Sparsectrl: Adding sparse controls to text-to-video diffusion models. In European Conference on Computer Vision . Springer, 330–348
Yuwei Guo, Ceyuan Yang, Anyi Rao, Maneesh Agrawala, Dahua Lin, and Bo Dai. 2025 · 2025
Closest in time.
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation. In European Conference on Computer Vision . Springer, 360–378
Shaowei Liu, Zhongzheng Ren, Saurabh Gupta, and Shenlong Wang. 2025 · 2025
Closest in time.
Draganything: Motion control for anything using entity representation. In European Conference on Computer Vision . Springer, 331–348
Weijia Wu, Zhuang Li, Yuchao Gu, Rui Zhao, Yefei He, David Junhao Zhang, Mike Zheng Shou, Yan Li, Tingting Gao, and Di Zhang. 2025 · 2025
Closest in time.
Dynamicrafter: Animating open-domain images with video diffusion priors. In European Conference on Computer Vision . Springer, 399–417
Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, and Tien-Tsin Wong. 2025 · 2025
Closest in time.
Physdreamer: Physics-based interaction with 3d objects via video generation. In European Conference on Computer Vision . Springer, 388–406
Tianyuan Zhang, Hong-Xing Yu, Rundi Wu, Brandon Y Feng, Changxi Zheng, Noah Snavely, Jiajun Wu, and William T Freeman. 2025 · 2025
Closest in time.
Reconstruction and simulation of elastic objects with spring-mass 3D Gaussians. In European Conference on Computer Vision . Springer, 407–423
Licheng Zhong, Hong-Xing Yu, Jiajun Wu, and Yunzhu Li. 2025 · 2025
Closest in time.
Fsgs: Real-time few-shot view synthesis using gaussian splatting. In European Conference on Computer Vision . Springer, 145–163
Zehao Zhu, Zhiwen Fan, Yifan Jiang, and Zhangyang Wang. 2025 · 2025
Closest in time.