Fetching the paper…
Reading the bibliography…
Today's driving world models can generate remarkably realistic dash-cam videos, yet no single model excels universally.
Quo vadis, action recognition? A new model and the kinetics dataset
Joao Carreira et al · 2017
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Thomas Unterthiner et al · 2018
Earlier work this paper cites.
nuScenes: A multimodal dataset for autonomous driving
Holger Caesar et al · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron et al · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy et al · 2021
Earlier work this paper cites.
TransReID: Transformer-based object re-identification
Shuting He et al · 2021
Earlier work this paper cites.
MUSIQ: Multi-scale image quality transformer
Junjie Ke et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford et al · 2021
Earlier work this paper cites.
LoFTR: Detector-free local feature matching with transformers
Jiaming Sun et al · 2021
Earlier work this paper cites.
SegFormer: Simple and efficient design for semantic segmentation with transformers
Enze Xie et al · 2021
Earlier work this paper cites.
MagicDrive: Street view generation with diverse 3D geometry control
Ruiyuan Gao et al · 2023
Earlier work this paper cites.
Planning-oriented autonomous driving
Yihan Hu et al · 2023
Earlier work this paper cites.
VAD: Vectorized scene representation for efficient autonomous driving
Bo Jiang et al · 2023
Cited alongside, same era.
3D Gaussian splatting for real-time radiance field rendering
Bernhard Kerbl et al · 2023
Cited alongside, same era.
BEVFusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation
Zhijian Liu et al · 2023
Cited alongside, same era.
NAVSIM: Data-driven non-reactive autonomous vehicle simulation and benchmarking
Daniel Dauner et al · 2024
Cited alongside, same era.
ADA-Track: End-to-end multi-camera 3D multi-object tracking with alternating detection and association
Shuxiao Ding et al · 2024
Cited alongside, same era.
VBench: Comprehensive benchmark suite for video generative models
Ziqi Huang et al · 2024
Shuai Bai et al · 2025
Later among the works it cites.
Genie 3: A new frontier for world models, 2025
Philip J. Ball et al · 2025
Later among the works it cites.
DiST-4D: Disentangled spatiotemporal diffusion with metric depth for 4D driving scene generation
Jiazhe Guo et al · 2025
Later among the works it cites.
3D and 4D world modeling: A survey
Lingdong Kong et al · 2025
Later among the works it cites.
SAM 2: Segment anything in images and videos
Nikhila Ravi et al · 2025
Later among the works it cites.
Gen3C: 3D-informed world-consistent video generation with precise camera control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
DreamForge: Motion-aware autoregressive video generation for multi-view driving scenes
Jianbiao Mei et al · 2024
Cited alongside, same era.
Genie 2: A large-scale foundation world model, 2024
Jack Parker-Holder et al · 2024
Cited alongside, same era.
Grounded SAM: Assembling open-world models for diverse visual tasks
Tianhe Ren et al · 2024
Cited alongside, same era.
SparseOCC: Rethinking sparse latent representation for vision-based semantic occupancy prediction
Pin Tang et al · 2024
Cited alongside, same era.
Depth anything v2
Lihe Yang et al · 2024
Cited alongside, same era.
Cross-video identity correlating for person re-identification pre-training
Jialong Zuo et al · 2024
Cited alongside, same era.
Xuanchi Ren et al · 2025
Later among the works it cites.
GAIA-2: A controllable multi-view generative world model for autonomous driving
Lloyd Russell et al · 2025
Later among the works it cites.
Open Driving World Models (OpenDWM)
SenseTime-FVG · 2025
Later among the works it cites.
RLGF: Reinforcement learning with geometric feedback for autonomous driving video generation
Tianyi Yan et al · 2025
Later among the works it cites.
DriveDreamer-2: LLM-enhanced world models for diverse driving video generation
Guosheng Zhao et al · 2025
Later among the works it cites.
WorldLens: Full-spectrum evaluations of driving world models in real world
Ao Liang et al · 2026
Closest in time.
OneVL: One-step latent reasoning and planning with vision-language explanation
Jinghui Lu et al · 2026
Closest in time.