Fetching the paper…
Reading the bibliography…
Human drivers adeptly navigate complex scenarios by utilizing rich attentional semantics, but the current autonomous systems struggle to replicate this ability, as they often lose critical semantic information when converting 2D observations into 3D space.
1905
Earlier work this paper cites.
D. Badre, “Cognitive control, hierarchy, and the rostro–caudal organization of the frontal lobes,” Trends in Cognitive Sciences
2008
Earlier work this paper cites.
M. Werling, J. Ziegler, S. Kammel, and S. Thrun, “Optimal trajectory generation for dynamic street scenarios in a fren x00e9; t frame,” in 2010 IEEE International Conference on Robotics and Automation
2010
Earlier work this paper cites.
A. Vaswani, “Attention is All you Need,” Advances in Neural Information Processing Systems
2017
Earlier work this paper cites.
W. Zeng, W. Luo, S. Suo, A. Sadat, B. Yang, S. Casas, and R. Urtasun, “End-To-End Interpretable Neural Motion Planner,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2019
Earlier work this paper cites.
A. Sadat, M. Ren, A. Pokrovsky, Y.-C. Lin, E. Yumer, and R. Urtasun, “Jointly Learnable Behavior and Trajectory Planning for Self-Driving Vehicles,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems
2019
Earlier work this paper cites.
C. Lu, M. J. G. Van De Molengraft, and G. Dubbelman, “Monocular Semantic Occupancy Grid Mapping With Convolutional Variational Encoder–Decoder Networks,” IEEE Robotics and Automation Letters
2019
Earlier work this paper cites.
F. Codevilla, E. Santana, A. M. López, and A. Gaidon, “Exploring the Limitations of Behavior Cloning for Autonomous Driving,” in Proceedings of the IEEE/CVF International Conference on Computer Vision
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, volume 1 (long and short papers)
2019
Earlier work this paper cites.
A. Sadat, S. Casas, M. Ren, X. Wu, P. Dhawan, and R. Urtasun, “Perceive, Predict, and Plan: Safe Motion Planning Through Interpretable Semantic Representations,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16
2020
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuScenes: A Multimodal Dataset for Autonomous Driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
Earlier work this paper cites.
B. Pan, J. Sun, H. Y. T. Leung, A. Andonian, and B. Zhou, “Cross-View Semantic Segmentation for Sensing Surroundings,” IEEE Robotics and Automation Letters
2020
Earlier work this paper cites.
T. Roddick and R. Cipolla, “Predicting Semantic Map Representations From Images Using Pyramid Occupancy Networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
Earlier work this paper cites.
J. Philion and S. Fidler, “Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIV 16
2020
Earlier work this paper cites.
D. Chen, B. Zhou, V. Koltun, and P. Krähenbühl, “Learning by Cheating,” in Conference on Robot Learning
2020
Earlier work this paper cites.
A. Prakash, K. Chitta, and A. Geiger, “Multi-Modal Fusion Transformer for End-to-End Autonomous Driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al
2021
Cited alongside, same era.
S. Casas, A. Sadat, and R. Urtasun, “MP3: A Unified Model To Map, Perceive, Predict and Plan,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2021
Cited alongside, same era.
H. Wang, P. Cai, Y. Sun, L. Wang, and M. Liu, “Learning Interpretable End-to-End Vision-Based Motion Planning for Autonomous Driving with Optical Flow Distillation,” in 2021 IEEE International Conference on Robotics and Automation
2021
Cited alongside, same era.
A. Hu, Z. Murez, N. Mohan, S. Dudas, J. Hawke, V. Badrinarayanan, R. Cipolla, and A. Kendall, “FIERY: Future Instance Prediction in Bird’s-Eye View From Surround Monocular Cameras,” in Proceedings of the IEEE/CVF International Conference on Computer Vision
2021
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Li, Z. Yu, S. Lan, J. Li, J. Kautz, T. Lu, and J. M. Alvarez, “Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2024
Later among the works it cites.
L. Chen, O. Sinavski, J. Hünermann, A. Karnsund, A. J. Willmott, D. Birch, D. Maund, and J. Shotton, “Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving,” in 2024 IEEE International Conference on Robotics and Automation
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Wang, V. C. Guizilini, T. Zhang, Y. Wang, H. Zhao, and J. Solomon, “DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries,” in Conference on Robot Learning
2022
Cited alongside, same era.
Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, et al
2022
Cited alongside, same era.
K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu, et al
2022
Cited alongside, same era.
S. Hu, L. Chen, P. Wu, H. Li, J. Yan, and D. Tao, “ST-P3: End-to-End Vision-Based Autonomous Driving via Spatial-Temporal Feature Learning,” in European Conference on Computer Vision
2022
Cited alongside, same era.
B. Jiang, S. Chen, Q. Xu, B. Liao, J. Chen, H. Zhou, Q. Zhang, W. Liu, C. Huang, and X. Wang, “VAD: Vectorized Scene Representation for Efficient Autonomous Driving,” in Proceedings of the IEEE/CVF International Conference on Computer Vision
2023
Cited alongside, same era.
Y. Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang, et al
2023
Cited alongside, same era.
J. Gu, C. Hu, T. Zhang, X. Chen, Y. Wang, Y. Wang, and H. Zhao, “ViP3D: End-to-End Visual Trajectory Prediction via 3D Agent Queries,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2023
Cited alongside, same era.
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,” in International Conference on Machine Learning
2023
Cited alongside, same era.
Later among the works it cites.
Z. Xu, Y. Zhang, E. Xie, Z. Zhao, Y. Guo, K.-Y. K. Wong, Z. Li, and H. Zhao, “DriveGPT4: Interpretable End-to-End Autonomous Driving Via Large Language Model,” IEEE Robotics and Automation Letters
2024
Later among the works it cites.
Y. Zhou, L. Huang, Q. Bu, J. Zeng, T. Li, H. Qiu, H. Zhu, M. Guo, Y. Qiao, and H. Li, “Embodied Understanding of Driving Scenarios,” in European Conference on Computer Vision
2024
Later among the works it cites.
C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, and H. Li, “DriveLM: Driving with Graph Visual Question Answering,” in European Conference on Computer Vision
2024
Later among the works it cites.
T. Qian, J. Chen, L. Zhuo, Y. Jiao, and Y.-G. Jiang, “NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario,” in Proceedings of the AAAI Conference on Artificial Intelligence
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Moon, H. Woo, H. Park, H. Jung, R. Mahjourian, H.-g. Chi, H. Lim, S. Kim, and J. Kim, “VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions,” in European Conference on Computer Vision
2024
Later among the works it cites.
X. Yi, H. Xu, H. Zhang, L. Tang, and J. Ma, “Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2024
Later among the works it cites.
W. Zheng, R. Song, X. Guo, C. Zhang, and L. Chen, “GenAD: Generative End-to-End Autonomous Driving,” in European Conference on Computer Vision
2024
Later among the works it cites.
W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, S. XiXuan, et al
2025
Closest in time.