Fetching the paper…
Reading the bibliography…
World modeling has become a cornerstone in AI research, enabling agents to understand, represent, and predict the dynamic environments they inhabit.
Seeing through fog without seeing fog: Deep multimodal sensor fusion in unseen adverse weather
Mario Bijelic et al · 1902
Earlier work this paper cites.
nuScenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, et al · 1903
Earlier work this paper cites.
Zheng Tang, Milind Naphade, Ming-Yu Liu, et al · 1903
Earlier work this paper cites.
SemanticKITTI: A dataset for semantic scene understanding of LiDAR sequences
Jens Behley, Martin Garbade, Andres Milioto, et al · 1904
Earlier work this paper cites.
3D point cloud generative adversarial network based on tree structured graph convolutions
Dong Wook Shu, Sung Woo Park, and Junseok Kwon · 1905
Earlier work this paper cites.
3D multi-object tracking: A baseline and new evaluation metrics
Xinshuo Weng et al · 1907
Earlier work this paper cites.
Scalability in perception for autonomous driving: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, et al · 1912
Earlier work this paper cites.
Yohann Cabon et al · 2001
Earlier work this paper cites.
NeRF: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2003
Earlier work this paper cites.
Convolutional occupancy networks
Songyou Peng et al · 2003
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang et al · 2004
Earlier work this paper cites.
Multiple object tracking performance metrics and evaluation in a smart room environment
Keni Bernardin, Alexander Elbs, and Rainer Stiefelhagen · 2006
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho et al · 2006
Earlier work this paper cites.
One thousand and one hours: Self-driving motion prediction dataset
John Houston, Guido Zuidhof, Luca Bergamini, et al · 2006
Earlier work this paper cites.
Scope of validity of PSNR in image/video quality assessment
Quan Huynh-Thu and Mohammed Ghanbari · 2008
Earlier work this paper cites.
Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3D
Jonah Philion and Sanja Fidler · 2008
Earlier work this paper cites.
ShapeAssembly: Learning to generate programs for 3D shape structure synthesis
R Kenny Jones, Theresa Barton, Xianghao Xu, et al · 2009
Earlier work this paper cites.
The Pascal visual object classes (VOC) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2010
Earlier work this paper cites.
SMARTS: Scalable multi-agent reinforcement learning training school for autonomous driving
Ming Zhou, Jun Luo, Julian Villella, et al · 2010
Earlier work this paper cites.
Are we ready for autonomous driving? the KITTI vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
iGibson 1.0: A simulation environment for interactive tasks in large realistic scenes
Bokui Shen, Fei Xia, Chengshu Li, et al · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from RGBD images
Nathan Silberman et al · 2012
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma, Max Welling, et al · 2013
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin et al · 2014
Earlier work this paper cites.
Improved techniques for training GANs
Tim Salimans et al · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
CARLA: An open urban driving simulator
Alexey Dosovitskiy et al · 2017
Earlier work this paper cites.
GANs trained by a two-time-scale update rule converge to a local Nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, et al · 2017
Earlier work this paper cites.
PointNet: Deep learning on point sets for 3D classification and segmentation
Charles R Qi et al · 2017
Earlier work this paper cites.
Semantic scene completion from a single depth image
Shuran Song, Fisher Yu, Andy Zeng, et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord et al · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, et al · 2017
Earlier work this paper cites.
Mikołaj Bińkowski et al · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Thomas Unterthiner et al · 2018
Earlier work this paper cites.
3D reconstruction of disaster scenes for urban search and rescue
Styliani Verykokou et al · 2018
Earlier work this paper cites.
Gibson env: real-world perception for embodied agents
Fei Xia et al · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, et al · 2018
Earlier work this paper cites.
Stereo magnification: Learning view synthesis using multiplane images
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely · 2018
Earlier work this paper cites.
Occupancy networks: Learning 3D reconstruction in function space
Lars Mescheder, Michael Oechsle, Michael Niemeyer, et al · 2019
Earlier work this paper cites.
RangeNet++: Fast and accurate LiDAR semantic segmentation
Andres Milioto et al · 2019
Earlier work this paper cites.
Kaichun Mo et al · 2019
Earlier work this paper cites.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
Jacob Austin, Daniel D Johnson, Jonathan Ho, Daniel Tarlow, and Rianne Van Den Berg · 2021
Earlier work this paper cites.
Physion: Evaluating physical prediction from vision in humans and machines
Daniel M Bear et al · 2021
Earlier work this paper cites.
nuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles
Holger Caesar, Juraj Kabzan, Kok Seang Tan, et al · 2021
Earlier work this paper cites.
Manipulathor: A framework for visual object manipulation
Kiana Ehsani, Winson Han, Alvaro Herrasti, et al · 2021
Earlier work this paper cites.
Dynamic view synthesis from dynamic monocular video
Chen Gao et al · 2021
Earlier work this paper cites.
MUSIQ: Multi-scale image quality transformer
Junjie Ke et al · 2021
Earlier work this paper cites.
Infinite nature: Perpetual view generation of natural scenes from a single image
Andrew Liu, Richard Tucker, Varun Jampani, Ameesh Makadia, Noah Snavely, and Angjoo Kanazawa · 2021
Earlier work this paper cites.
Learning to drop points for LiDAR scan synthesis
Kazuto Nakashima and Ryo Kurazume · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, et al · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, et al · 2021
Earlier work this paper cites.
LoFTR: Detector-free local feature matching with transformers
Jiaming Sun, Zehong Shen, Yuang Wang, et al · 2021
Earlier work this paper cites.
Habitat 2.0: Training home assistants to rearrange their habitat
Andrew Szot, Alexander Clegg, Eric Undersander, et al · 2021
Earlier work this paper cites.
Argoverse 2: Next generation datasets for self-driving perception and forecasting
Benjamin Wilson, William Qi, Tanmay Agarwal, et al · 2021
Earlier work this paper cites.
PandaSet: Advanced sensor suite dataset for autonomous driving
Pengchuan Xiao, Zhenlei Shao, Steven Hao, et al · 2021
Earlier work this paper cites.
MonoScene: Monocular 3D semantic scene completion
Anh-Quan Cao and Raoul De Charette · 2022
Earlier work this paper cites.
MaskGIT: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, et al · 2022
Earlier work this paper cites.
Image-based traffic signal control via world models
Xingyuan Dai, Chen Zhao, Xiao Wang, et al · 2022
Earlier work this paper cites.
ProcTHOR: Large-scale embodied AI using procedural generation
Matt Deitke, Eli VanderBilt, Alvaro Herrasti, et al · 2022
Earlier work this paper cites.
Panoptic nuScenes: A large-scale benchmark for LiDAR panoptic segmentation and tracking
Whye Kit Fong, Rohit Mohan, Juana Valeria Hurtado, et al · 2022
Earlier work this paper cites.
Differentiable raycasting for self-supervised occupancy forecasting
Tarasha Khurana, Peiyun Hu, Achal Dave, et al · 2022
Earlier work this paper cites.
Diffusion deformable model for 4D temporal medical image generation
Boah Kim and Jong Chul Ye · 2022
Earlier work this paper cites.
Immersive neural graphics primitives
Ke Li et al · 2022
Earlier work this paper cites.
KITTI-360: A novel dataset and benchmarks for urban scene understanding in 2D and 3D
Yiyi Liao, Jun Xie, and Andreas Geiger · 2022
Earlier work this paper cites.
Capturing, reconstructing, and simulating: the UrbanScene3D dataset
Liqiang Lin, Yilin Liu, Yue Hu, et al · 2022
Earlier work this paper cites.
Flow matching for generative modeling
Yaron Lipman et al · 2022
Earlier work this paper cites.
Learning latent graph dynamics for visual manipulation of deformable objects
Xiao Ma, David Hsu, and Wee Sun Lee · 2022
Earlier work this paper cites.
Jaideep Pathak et al · 2022
Earlier work this paper cites.
DreamFusion: Text-to-3D using 2D diffusion
Ben Poole et al · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
MotionSC: Data set and network for real-time semantic mapping in dynamic environments
Joey Wilson, Jingyu Song, Yuewei Fu, et al · 2022
Earlier work this paper cites.
Runsheng Xu et al · 2022
Earlier work this paper cites.
Cross-view transformers for real-time map-view semantic segmentation
Brady Zhou and Philipp Krähenbühl · 2022
Earlier work this paper cites.
Learning to generate realistic LiDAR point clouds
Vlas Zyrianov et al · 2022
Earlier work this paper cites.
Self-supervised learning from images with a joint-embedding predictive architecture
Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas · 2023
Cited alongside, same era.
HyperReel: High-fidelity 6-DoF video with ray-conditioned sampling
Benjamin Attal, Jia-Bin Huang, Christian Richardt, et al · 2023
Cited alongside, same era.
Accurate medium-range global weather forecasting with 3D neural networks
Kaifeng Bi, Lingxi Xie, Hengheng Zhang, et al · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, et al · 2023
Cited alongside, same era.
4D simulation research in construction: A systematic mapping study
Conrad Boton et al · 2023
DIO: Decomposable implicit 4D occupancy-flow world model
Christopher Diehl, Quinlan Sykora, Ben Agro, et al · 2025
Closest in time.
Haotian Dong, Xin Wang, Di Lin, et al · 2025
Closest in time.
Yi Du et al · 2025
Closest in time.
WorldScore: A unified evaluation benchmark for world generation
Haoyi Duan et al · 2025
Closest in time.
FreeSim: Toward free-viewpoint camera simulation in driving scenes
Lue Fan, Hao Zhang, Qitai Wang, et al · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
AnchorFormer: Point cloud completion from discriminative nodes
Zhikai Chen et al · 2023
Cited alongside, same era.
MagicDrive: Street view generation with diverse 3D geometry control
Ruiyuan Gao et al · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Cited alongside, same era.
Lin Guan et al · 2023
Cited alongside, same era.
ADriver-I: A general world model for autonomous driving
Fan Jia, Weixin Mao, Yingfei Liu, et al · 2023
Cited alongside, same era.
3D gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis · 2023
Cited alongside, same era.
Point cloud forecasting as a proxy for 4D occupancy forecasting
Tarasha Khurana et al · 2023
Cited alongside, same era.
Closest in time.
A survey of world models for autonomous driving
Tuo Feng, Wenguan Wang, and Yi Yang · 2025
Closest in time.
Foundation models in robotics: Applications, challenges, and the future
Roya Firoozi, Johnathan Tucker, Stephen Tian, et al · 2025
Closest in time.
Embodied AI agents: Modeling the world
Pascale Fung, Yoram Bachrach, Asli Celikyilmaz, Kamalika Chaudhuri, Delong Chen, Willy Chung, Emmanuel Dupoux, Hongyu Gong, Hervé Jégou, Alessandro Lazaric, et al · 2025
Closest in time.
Unraveling the effects of synthetic data on end-to-end autonomous driving
Junhao Ge et al · 2025
Closest in time.
DiST-4D: Disentangled spatiotemporal diffusion with metric depth for 4D driving scene generation
Jiazhe Guo, Yikang Ding, Xiwu Chen, et al · 2025
Closest in time.
Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, et al · 2025
Closest in time.
3D4D: An interactive, editable, 4D world model via 3D video generation
Yunhong He, Zhengqing Yuan, Zhengzhong Tu, Yanfang Ye, and Lichao Sun · 2025
Closest in time.
LiDAR-EDIT: LiDAR data generation by editing the object layouts in real-world scenes
Shing-Hei Ho, Bao Thach, and Minghan Zhu · 2025
Closest in time.
CoGen: 3D consistent video generation via adaptive conditioning for autonomous driving
Yishen Ji, Ziyue Zhu, Zhenxin Zhu, et al · 2025
Closest in time.
MGVQ: Could VQ-VAE beat VAE? a generalizable tokenizer with multi-group quantization
Mingkai Jia et al · 2025
Closest in time.
How far is video generation from world model: A physical law perspective
Bingyi Kang et al · 2025
Closest in time.
LOGen: Toward LiDAR object generation by point diffusion
Ellington Kirby et al · 2025
Closest in time.
Multi-modal data-efficient 3D scene understanding for autonomous driving
Lingdong Kong, Xiang Xu, Jiawei Ren, et al · 2025
Closest in time.
I2-World: Intra-inter tokenization for efficient dynamic 4D scene forecasting
Zhimin Liao, Ping Wei, Ruijie Zhang, et al · 2025
Closest in time.
Exploring the evolution of physics cognition in video generation: A survey
Minghui Lin et al · 2025
Closest in time.
A survey: Learning embodied intelligence from physical simulators and world models
Xiaoxiao Long et al · 2025
Closest in time.
DreamDrive: Generative 4D scene modeling from street view images
Jiageng Mao, Boyi Li, Boris Ivanovic, et al · 2025
Closest in time.
LiDPM: Rethinking point diffusion for LiDAR scene completion
Tetiana Martyniuk et al · 2025
Closest in time.
Vision-centric 4D occupancy forecasting and planning via implicit residual world models
Jianbiao Mei, Yu Yang, Xuemeng Yang, Licheng Wen, Jiajun Lv, Botian Shi, and Yong Liu · 2025
Closest in time.
Dreamland: Controllable world creation with simulator and generative models
Sicheng Mo et al · 2025
Closest in time.
Fast LiDAR data generation with rectified flows
Kazuto Nakashima et al · 2025
Closest in time.
Towards generating realistic 3D semantic training data for autonomous driving
Lucas Nunes et al · 2025
Closest in time.
Four principles for physically interpretable world models
Jordan Peper et al · 2025
Closest in time.
Cosmos-Drive-Dreams: Scalable synthetic driving data generation with world foundation models
Xuanchi Ren, Yifan Lu, Tianshi Cao, et al · 2025
Closest in time.
GAIA-2: A controllable multi-view generative world model for autonomous driving
Lloyd Russell, Anthony Hu, Lorenzo Bertoni, et al · 2025
Closest in time.
Open driving world models (OpenDWM)
SenseTime-FVG · 2025
Closest in time.
SceneDiffuser++: City-scale traffic simulation via a generative world model
Shuhan Tan, John Lambert, Hong Jeon, et al · 2025
Closest in time.
Hunyuan-GameCraft-2: Instruction-following interactive game world model
Junshu Tang, Jiacheng Liu, Jiaqi Li, Longhuang Wu, Haoyu Yang, Penghao Zhao, Siruis Gong, Xiang Yuan, Shuai Shao, Linfeng Zhang, and Qinglin Lu · 2025
Closest in time.
PAN: A world model for general, interactable, and long-horizon world simulation
PAN Team et al · 2025
Closest in time.
The role of world models in shaping autonomous driving: A comprehensive survey
Sifan Tu et al · 2025
Closest in time.
ProphetDWM: A driving world model for rolling out future actions and videos
Xiaodong Wang and Peixi Peng · 2025
Closest in time.
LiDARDraft: Generating LiDAR point cloud from versatile inputs
Haiyun Wei, Fan Lu, Yunwei Zhu, Zehan Zheng, Weiyi Xue, Lin Shao, Xudong Zhang, Ya Wu, Rong Fu, and Guang Chen · 2025
Closest in time.
Beichen Wen et al · 2025
Closest in time.
SG-LDM: Semantic-guided lidar generation via latent-aligned diffusion
Zhengkang Xiang, Zizhao Li, Amir Khodabandeh, and Kourosh Khoshelham · 2025
Closest in time.
Assessing adaptive world models in machines with novel games
Lance Ying et al · 2025
Closest in time.
Uni-Gaussians: Unifying camera and LiDAR simulation with Gaussians for dynamic driving scenarios
Zikang Yuan et al · 2025
Closest in time.
Rethinking driving world model as synthetic data generator for perception tasks
Kai Zeng, Zhanqian Wu, Kaixin Xiong, Xiaobao Wei, Xiangyu Guo, Zhenxin Zhu, Kalok Ho, Lijun Zhou, Bohan Zeng, Ming Lu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, and Wentao Zhang · 2025
Closest in time.
MuDG: Taming multi-modal diffusion with Gaussian splatting for urban scene reconstruction
Yingshuang Zou et al · 2025
Closest in time.
GaussianWorld: Gaussian world model for streaming 3D occupancy prediction
Sicheng Zuo, Wenzhao Zheng, Yuanhui Huang, Jie Zhou, and Jiwen Lu · 2025
Closest in time.
PointDiffusion: Diffusion-based scene completion in the point cloud domain
Chidera Agbasiere et al · 2026
Closest in time.
EditSSC: Toward editable semantic occupancy scenes with unconditional diffusion models
Fatima Balde, Raoul de Charette, and Alexandre Boulch · 2026
Closest in time.
Adversarially guided diffusion for LiDAR range image synthesis
Stavros Bouras et al · 2026
Closest in time.
The seriality gap in video diffusion models
Jorge Diaz Chao et al · 2026
Closest in time.
GEM: Gaussian evolution model for occupancy forecasting and motion planning
Cheng Chen, Hao Huang, and Saurabh Bagchi · 2026
Closest in time.
OWMDrive: Causality-aware end-to-end autonomous driving via 4D occupancy world model
Junjie Cheng et al · 2026
Closest in time.
Agentic world modeling: Foundations, capabilities, laws, and beyond
Meng Chu, Xuan Billy Zhang, Kevin Qinghong Lin, Lingdong Kong, Jize Zhang, Teng Tu, Weijian Ma, Ziqi Huang, Senqiao Yang, Wei Huang, Yeying Jin, Zhefan Rao, Jinhui Ye, Xinyu Lin, Xichen Zhang, Qisheng Hu, Shuai Yang, Leyang Shen, Wei Chow, Yifei Dong, Fengyi Wu, Quanyu Long, Bin Xia, Shaozuo Yu, Mingkang Zhu, Wenhu Zhang, Jiehui Huang, Haokun Gui, Runyi Li, Chenyu Tang, Dong Huang, Xuhang Chen, Rui Liu, Chengzu Li, Shiyi Du, Xu Huang, Haoxuan Che, Long Chen, Qifeng Chen, Wenya Wang, Wenxuan Zhang, Xiaojuan Qi, Yang Deng, Yanwei Li, Mike Zheng Shou, Zhi-Qi Cheng, See-Kiong Ng, Ziwei Liu, Philip Torr, and Jiaya Jia · 2026
Closest in time.
Glob3R: Global structure-from-motion with 3D foundation models
Junyuan Deng et al · 2026
Closest in time.
Zeying Gong, Yangyi Zhong, Yiyi Ding, Tianshuai Hu, Guoyang Zhao, Lingdong Kong, Rong Li, Jiadi You, and Junwei Liang · 2026
Closest in time.
CascadeOcc: Rethinking 3D occupancy world models with cascaded VQ representations
Kyumin Hwang et al · 2026
Closest in time.
Hy-world 2.0: A multi-modal world model for reconstructing, generating, and simulating 3D worlds
Team HY-World et al · 2026
Closest in time.
INSPATIO-WORLD: A real-time 4D world simulator via spatiotemporal autoregressive modeling
InSpatio Team, Donghui Shen, Guofeng Zhang, Haomin Liu, Haoyu Ji, Hujun Bao, Hongjia Zhai, Jialin Liu, Jing Guo, Nan Wang, Siji Pan, Weihong Pan, Weijian Xie, Xianbin Liu, Xiaojun Xiang, Xiaoyu Zhang, Xinyu Chen, Yifu Wang, Yipeng Chen, Zhenzhou Fan, Zhewen Le, Zhichao Ye, and Ziqiang Zhao · 2026
Closest in time.
Cam2Sim: Neural scenario reconstruction for closed-loop autonomous driving simulation
Davide Jannussi et al · 2026
Closest in time.
Is your driving world model an all-around player?
Lingdong Kong, Ao Liang, Tianyi Yan, Hongsi Liu, Wesley Yang, Ziqi Huang, Xian Sun, Wei Yin, Jialong Zuo, Yixuan Hu, Dekai Zhu, Dongyue Lu, Youquan Liu, Guangfeng Jiang, Linfeng Li, Xiangtai Li, Long Zhuo, Lai Xing Ng, Benoit R. Cottereau, Changxin Gao, Liang Pan, Wei Tsang Ooi, and Ziwei Liu · 2026
Closest in time.
Xiaomi OneVL: One-step latent reasoning and planning with vision-language explanation
Jinghui Lu, Jiayi Guan, Zhijian Huang, Jinlong Li, Guang Li, Lingdong Kong, Yingyan Li, Han Wang, Shaoqing Xu, Yuechen Luo, Fang Li, Chenxu Dang, Junli Wang, Tao Xu, Jing Wu, Jianhua Wu, Xiaoshuai Hao, Wen Zhang, Tianyi Jiang, Lingfeng Zhang, Lei Zhou, Yingbo Tang, Jie Wang, Yinfeng Gao, Xizhou Bu, Haochen Tian, Yihang Qiu, Feiyang Jia, Lin Liu, Yigu Ge, Hanbing Li, Yuannan Shen, Jianwei Cui, Hongwei Xie, Bing Wang, Haiyang Sun, Jingwei Zhao, Jiahui Huang, Pei Liu, Zeyu Zhu, Yuncheng Jiang, Zibin Guo, Chuhong Gong, Hanchao Leng, Kun Ma, Naiyan Wang, Guang Chen, Kuiyuan Yang, Hangjun Ye, and Long Chen · 2026
Closest in time.
Zijie Meng, Yufei Liu, Chengqian Ma, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Shuqin Chen, Weichen Xu, Jiquan Yuan, and Miao Zhang · 2026
Closest in time.
ForecastOcc: Vision-based semantic occupancy forecasting
Riya Mohan, Juana Valeria Hurtado, Rohit Mohan, and Abhinav Valada · 2026
Closest in time.
Validate the dream before you trust its verdict: Admissibility for world-model simulators
Christian Oefinger et al · 2026
Closest in time.
AutoWorld: Scaling multi-agent traffic simulation with self-supervised world models
Mozhgan Pourkeshavatz, Tianran Liu, and Nicholas Rhinehart · 2026
Closest in time.
4DLidarOpen: An open 4D FMCW lidar dataset for motion-aware autonomous driving
Kane Qian, Xin Zhao, Yining Shi, Rujun Yan, Zhengqing Pan, Kaojin Zhu, Mengmeng Yang, Kai Sun, Diange Yang, and Kun Jiang · 2026
Closest in time.
Wentao Qu et al · 2026
Closest in time.
Solaris: Building a multiplayer video world model in minecraft
Georgy Savva, Oscar Michel, Daohan Lu, Suppakit Waiwitlikhit, Timothy Meehan, Dhairya Mishra, Srivats Poddar, Jack Lu, and Saining Xie · 2026
Closest in time.
R3DPA: Leveraging 3D representation alignment and rgb pretrained priors for lidar scene generation
Nicolas Sereyjol-Garros, Ellington Kirby, Victor Besnier, and Nermin Samet · 2026
Closest in time.
Lyra 2.0: Explorable generative 3D worlds
Tianchang Shen, Sherwin Bahmani, Kai He, Sangeetha Grama Srinivasan, Tianshi Cao, Jiawei Ren, Ruilong Li, Zian Wang, Nicholas Sharp, Zan Gojcic, Sanja Fidler, Jiahui Huang, Huan Ling, Jun Gao, and Xuanchi Ren · 2026
Closest in time.
NoDrift3R: Raymap-guided coupling for drift-robust unposed feed-forward 3D reconstruction
Xiangyu Sun et al · 2026
Closest in time.
PrITTI: Primitive-based generation of controllable and editable 3D semantic urban scenes
Christina Ourania Tze, Daniel Dauner, Yiyi Liao, et al · 2026
Closest in time.
Matrix-game 3.0: Real-time and streaming interactive world model with long-horizon memory
Zile Wang, Zexiang Liu, Jiaxing Li, Kaichen Huang, Baixin Xu, Fei Kang, Mengyin An, Peiyu Wang, Biao Jiang, Yichen Wei, Yidan Xietian, Jiangbo Pei, Liang Hu, Boyi Jiang, Hua Xue, Zidong Wang, Haofeng Sun, Wei Li, Wanli Ouyang, Xianglong He, Yang Liu, Yangguang Li, and Yahui Zhou · 2026
Closest in time.
Songbur Wong et al · 2026
Closest in time.
VISA: VLM-guided instance semantic auditing for 3D occupancy world models
Ruiqi Xian et al · 2026
Closest in time.
PerpetualWonder: Long-horizon action-conditioned 4D scene generation
Jiahao Zhan, Zizhang Li, Hong-Xing Yu, and Jiajun Wu · 2026
Closest in time.
CP4D: Compositional physics-aware 4D scene generation
Hanxin Zhu, Cong Wang, Tianyu He, Long Chen, Xin Jin, Chen Gao, and Zhibo Chen · 2026
Closest in time.
Arif Hassan Zidan, Yi Pan, Hanqi Jiang, Ruiyu Yan, Wei Ruan, Zihao Wu, Lifeng Chen, Weihang You, Xinliang Li, Bowen Chen, Huawen Hu, Peilong Wang, Sizhuang Liu, Jing Zhang, Siyuan Li, Zhengliang Liu, Yu Bao, Lin Zhao, Lichao Sun, Dajiang Zhu, Xiang Li, Jinglei Lv, Quanzheng Li, Wei Liu, Tianming Liu, and Wei Zhang · 2026
Closest in time.