Fetching the paper…
Reading the bibliography…
World models for autonomous driving have the potential to dramatically improve the reasoning capabilities of today's systems.
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,” International Journal of Robotics Research (IJRR) , 2013
2013
Earlier work this paper cites.
K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation,” in Conf. on Empirical Methods in Natural Language Processing (EMNLP) , 2014
2014
Earlier work this paper cites.
B. Li, T. Zhang, and T. Xia, “Vehicle Detection from 3D Lidar Using Fully Convolutional Network,” in Rob. Sci. Sys. , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in ECCV , 2016
2016
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural Discrete Representation Learning,” in NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
S. Song, F. Yu, A. Zeng, A. X. Chang, M. Savva, and T. Funkhouser, “Semantic Scene Completion from a Single Depth Image,” in CVPR , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit et al. , “Attention is All you Need,” in NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “CARLA: An open urban driving simulator,” in CoRL , 2017
2017
Earlier work this paper cites.
D. Ha and J. Schmidhuber, “Recurrent world models facilitate policy evolution,” in NeurIPS , 2018
2018
Earlier work this paper cites.
A. H. Lang, S. Vora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “PointPillars: Fast Encoders for Object Detection From Point Clouds,” in CVPR , 2019
2019
Earlier work this paper cites.
X. Yan, J. Gao, J. Li, R. Zhang, Z. Li, R. Huang, and S. Cui, “Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene Completion,” in AAAI , 2020
2020
Earlier work this paper cites.
J. Philion and S. Fidler, “Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D,” in ECCV , 2020
2020
Earlier work this paper cites.
Waymo, “Introducing the 5th-generation waymo driver,” 2020. [Online]. Available: https://waymo.com/blog/2020/03/introducing-5th-generation-waymo-driver.html
2020
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn et al. , “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in ICLR , 2020
2020
Earlier work this paper cites.
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to Control: Learning Behaviors by Latent Imagination,” in ICLR , 2020
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh et al. , “Learning Transferable Visual Models From Natural Language Supervision,” in ICML , 2021
2021
Earlier work this paper cites.
S. W. Kim, , J. Philion, A. Torralba, and S. Fidler, “DriveGAN: Towards a Controllable High-Quality Neural Simulation,” in CVPR , 2021
2021
Earlier work this paper cites.
D. Hafner, T. Lillicrap, M. Norouzi, and J. Ba, “Mastering Atari with Discrete World Models,” in ICLR , 2021
2021
Earlier work this paper cites.
L. Chen, K. Lu, A. Rajeswaran, K. Lee et al. , “Decision Transformer: Reinforcement Learning via Sequence Modeling,” in NeurIPS , 2021
2021
Earlier work this paper cites.
M. Janner, Q. Li, and S. Levine, “Offline Reinforcement Learning as One Big Sequence Modeling Problem,” in NeurIPS , 2021
2021
Earlier work this paper cites.
C. B. Rist, D. Emmerichs, M. Enzweiler, and D. M. Gavrila, “Semantic Scene Completion using Local Deep Implicit Functions on LiDAR Data,” in IEEE TPAMI , 2021
2021
Earlier work this paper cites.
Z. Zhang, A. Liniger, D. Dai, F. Yu, and L. Van Gool, “End-to-end urban driving by imitating a reinforcement learning coach,” in ICCV , 2021
2021
Earlier work this paper cites.
K. Chitta, A. Prakash, B. Jaeger, Z. Yu, K. Renz, and A. Geiger, “TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving,” IEEE TPAMI , 2022
2022
Earlier work this paper cites.
Y. LeCun, “A Path Towards Autonomous Machine Intelligence,” OpenReview:BZ5a1r-kVsf , 2022
2022
Earlier work this paper cites.
A. Hu, G. Corrado, N. Griffiths, Z. Murez et al. , “Model-Based Imitation Learning for Urban Driving,” in NeurIPS , 2022
2022
Earlier work this paper cites.
Z. Gao, Y. Mu, R. Shen, C. Chen et al. , “Enhance Sample Efficiency and Robustness of End-to-end Urban Autonomous Driving via Semantic Masked World Model,” in NeurIPSW , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
P. Wu, A. Escontrela, D. Hafner, K. Goldberg, and P. Abbeel, “Daydreamer: World models for physical robot learning,” CoRL , 2022
2022
Earlier work this paper cites.
V. Micheli, E. Alonso, and F. Fleuret, “Transformers are Sample-Efficient World Models,” in ICLR , 2022
2022
Cited alongside, same era.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo et al. , “A Generalist Agent,” in TMLR , 2022
2022
Cited alongside, same era.
A.-Q. Cao and R. De Charette, “MonoScene: Monocular 3D Semantic Scene Completion,” in CVPR , 2022
2022
Cited alongside, same era.
J. Chen, S. E. Li, and M. Tomizuka, “Interpretable End-to-End Urban Autonomous Driving With Latent Deep Reinforcement Learning,” T-ITS , vol. 23, 2022
2022
Cited alongside, same era.
H. Shao, L. Wang, R. Chen, H. Li, and Y. Liu, “Safety-Enhanced Autonomous Driving Using Interpretable Sensor Fusion Transformer,” in CoRL , 2022
2022
Cited alongside, same era.
A. Carlson, M. S. Ramanagopal, N. Tseng, M. Johnson-Roberson, R. Vasudevan, and K. A. Skinner, “CLONeR: Camera-Lidar Fusion for Occupancy Grid-aided Neural Representations,” Rob. Aut. Lett. , vol. 8, no. 5, 2023
2023
Closest in time.
C. Sima, W. Tong, T. Wang, L. Chen et al. , “Scene as occupancy,” in ICCV , 2023
2023
Closest in time.
O. Contributors, “Openscene: The largest up-to-date 3d occupancy prediction benchmark in autonomous driving,” 2023. [Online]. Available: https://github.com/OpenDriveLab/OpenScene
2023
Closest in time.
T. Khurana, P. Hu, D. Held, and D. Ramanan, “Point Cloud Forecasting as a Proxy for 4D Occupancy Forecasting,” in CVPR , 2023
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Liang, H. Xie, K. Yu, Z. Xia et al. , “BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework,” in NeurIPS , 2022
2022
Cited alongside, same era.
S. Mehta and M. Rastegari, “Separable self-attention for mobile vision transformers,” TMLR , 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
A. Elluswamy, “Foundation Models for Autonomy,” 2023, CVPR Workshop on Autonomous Driving. [Online]. Available: https://www.youtube.com/watch?v=6x-Xb_uT7ts
2023
Cited alongside, same era.
Z. Liu, H. Tang, A. Amini, X. Yang, H. Mao, D. L. Rus, and S. Han, “BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation,” in ICRA , 2023
2023
Cited alongside, same era.
D. Bogdoll, L. Bosch, T. Joseph, H. Gremmelmaier, Y. Yang, and J. M. Zöllner, “Exploring the Potential of World Models for Anomaly Detection in Autonomous Driving,” in SSCI , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder et al. , “Genie: Generative interactive environments,” arXiv: 2402.15391 , 2024
2024
Closest in time.
J. Parker-Holder, P. Ball, J. Bruce, V. Dasagi et al. , “Genie 2: A large-scale foundation world model,” , 2024
2024
Closest in time.
2024
Closest in time.
A. Popov, A. Degirmenci, D. Wehr, S. Hegde et al. , “Mitigating covariate shift in imitation learning for autonomous vehicles using latent space generative world models,” arXiv 2409.16663 , 2024
2024
Closest in time.
S. Gao, J. Yang, L. Chen, K. Chitta et al. , “Vista: A generalizable driving world model with high fidelity and versatile controllability,” in NeurIPS , 2024
2024
Closest in time.
J. Yang, S. Gao, Y. Qiu, L. Chen et al. , “Generalized Predictive Model for Autonomous Driving,” in CVPR , 2024
2024
Closest in time.
Y. Wang, J. He, L. Fan, H. Li, Y. Chen, and Z. Zhang, “Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving,” in CVPR , 2024
2024
Closest in time.
L. Zhang, Y. Xiong, Z. Yang, S. Casas, R. Hu, and R. Urtasun, “Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion,” in ICLR , 2024
2024
Closest in time.
R. Chen, H. Park, B. Zhang, W. Shao, P. Luo, and A. Wong, “Trend: Unsupervised 3d representation learning via temporal forecasting for lidar perception,” arXiv: 2412.03054 , 2024
2024
Closest in time.
V. Zyrianov, H. Che, Z. Liu, and S. Wang, “Lidardm: Generative lidar simulation in a generated world,” arXiv 2404.02903 , 2024
2024
Closest in time.
Y. Zhang, S. Gong, K. Xiong, X. Ye et al. , “Bevworld: A multimodal world model for autonomous driving via unified bev latent space,” arXiv 2407.05679 , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Zhang, Y. Zhai, J. Mei, and Y. Hu, “Fusionocc: Multi-modal fusion for 3d occupancy prediction,” in ACM Multimedia 2024 , 2024
2024
Closest in time.
H. Jiang, T. Cheng, N. Gao, H. Zhang, W. Liu, and X. Wang, “Symphonize 3D Semantic Scene Completion with Contextual Instance Queries,” in CVPR , 2024
2024
Closest in time.
A. Hayler, F. Wimbauer, D. Muhle, C. Rupprecht, and D. Cremers, “S4C: Self-Supervised Semantic Scene Completion with Neural Fields,” in International Conference on 3D Vision (3DV) , 2024
2024
Closest in time.
2024
Closest in time.
Z. Yang, L. Chen, Y. Sun, and H. Li, “Visual point cloud forecasting enables scalable autonomous driving,” in CVPR , 2024
2024
Closest in time.
N. Agarwal, A. Ali, M. Bala, Y. Balaji et al. , “Cosmos world foundation model platform for physical ai,” arXiv: 2501.03575 , 2025
2025
Closest in time.
2025
Closest in time.
Y. Yang, J. Mei, Y. Ma, S. Du et al. , “Driving in the occupancy world: Vision-centric 4d occupancy forecasting and planning via world models for autonomous driving,” AAAI , 2025
2025
Closest in time.
H. Wang, X. Ye, F. Tao, C. Pan et al. , “Adawm: Adaptive world model based planning for autonomous driving,” in ICLR , 2025
2025
Closest in time.