Fetching the paper…
Reading the bibliography…
Evaluating vision-language models (VLMs) in urban driving contexts remains challenging, as existing benchmarks rely on open-ended responses that are ambiguous, annotation-intensive, and inconsistent to score.
Papineni, K., Roukos, S., Ward, T., and Zhu, W. J. (2002, July). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL)
2002
Earlier work this paper cites.
Lin, C. Y. (2004, July). ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out
2004
Earlier work this paper cites.
Banerjee, S. and Lavie, A. (2005, June). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization
2005
Earlier work this paper cites.
Vedantam, R., Lawrence Zitnick, C., and Parikh, D. (2015). CIDEr: Consensus-based image description evaluation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2015
Earlier work this paper cites.
Anderson, P., Fernando, B., Johnson, M., and Gould, S. (2016). SPICE: Semantic propositional image caption evaluation. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part V
2016
Earlier work this paper cites.
Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. (2017, October). CARLA: An open urban driving simulator. In Proceedings of the Conference on Robot Learning (CoRL)
2017
Earlier work this paper cites.
Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., and Beijbom, O. (2020). nuScenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Lu, J., Ye, X., Ren, Y., and Yang, Y. (2022). Good, better, best: Textual distractors generation for multiple-choice visual question answering via reinforcement learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., and Wei, F. (2023). Image as a foreign language: BEiT pretraining for vision and vision-language tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
OpenAI. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774
2024
Later among the works it cites.
Fu, D., Lei, W., Wen, L., Cai, P., Mao, S., Dou, M., and Qiao, Y. (2024, June). Limsim++: A closed-loop platform for deploying multimodal LLMs in autonomous driving. In Proceedings of the 2024 IEEE Intelligent Vehicles Symposium (IV)
2024
Later among the works it cites.
Cao, X., Zhou, T., Ma, Y., Ye, W., Cui, C., Tang, K., and Zheng, C. (2024). MapLM: A real-world large-scale vision-language benchmark for map and traffic scene understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2024
Later among the works it cites.
Qian, T., Chen, J., Zhuo, L., Jiao, Y., and Jiang, Y. G. (2024, March). NuScenes-QA: A multi-modal visual question answering benchmark for autonomous driving scenario. In Proceedings of the AAAI Conference on Artificial Intelligence
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Chen, L., Sinavski, O., Hünermann, J., Karnsund, A., Willmott, A. J., Birch, D., and Shotton, J. (2024, May). Driving with LLMs: Fusing object-level vector modality for explainable autonomous driving. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Khalili, B. and Smyth, A. W. (2024). SOD-YOLOv8—enhancing YOLOv8 for small object detection in aerial imagery and traffic scenes. Sensors
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Sima, C., Renz, K., Chitta, K., Chen, L., Zhang, H., Xie, C., and Li, H. (2024, September). DriveLM: Driving with graph visual question answering. In Proceedings of the European Conference on Computer Vision (ECCV)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Marcu, A. M., Chen, L., Hünermann, J., Karnsund, A., Hanotte, B., Chidananda, P., and Sinavski, O. (2024, September). LingoQA: Visual question answering for autonomous driving. In Proceedings of the European Conference on Computer Vision (ECCV)
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Ding, X., Han, J., Xu, H., Liang, X., Zhang, W., and Li, X. (2024). Holistic autonomous driving understanding by bird’s-eye-view injected multi-modal large models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2024
Later among the works it cites.
2024
Later among the works it cites.
Luo, H., Deng, Y., Shen, Y., Ng, S. K., and Chua, T. S. (2024). Chain-of-exemplar: Enhancing distractor generation for multimodal educational question generation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL)
2024
Later among the works it cites.
2025
Closest in time.