Fetching the paper…
Reading the bibliography…
Multimodal large language models (MLLMs) have demonstrated significant progress in semantic scene understanding and text-image alignment, with reasoning variants enhancing performance on more complex tasks involving mathematics and logic.
Im2gps: estimating geographic information from a single image
James Hays and Alexei A Efros · 2008
Earlier work this paper cites.
Learning from maps: Visual common sense for autonomous driving
Ari Seff and Jianxiong Xiao · 2016
Earlier work this paper cites.
Geostyle: Discovering fashion trends and events
Utkarsh Mall, Kevin Matzen, Bharath Hariharan, Noah Snavely, and Kavita Bala · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
A survey of embodied ai: From simulators to research tasks
Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
A review of high-definition map creation methods for autonomous driving
Zhibin Bao, Sabir Hossain, Haoxiang Lang, and Xianke Lin · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Earlier work this paper cites.
Geo-bench: Toward foundation models for earth monitoring
Alexandre Lacoste, Nils Lehmann, Pau Rodriguez, Evan Sherwin, Hannah Kerner, Björn Lütjens, Jeremy Irvin, David Dao, Hamed Alemohammad, Alexandre Drouin, et al · 2023
Earlier work this paper cites.
Kosmos-2: Grounding multimodal large language models to the world
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei · 2023
Earlier work this paper cites.
Lidar2map: In defense of lidar-based semantic map construction using online camera distillation
Song Wang, Wentong Li, Wenyu Liu, Xiaolu Liu, and Jianke Zhu · 2023
Earlier work this paper cites.
What you see is what you read? improving text-image alignment evaluation
Michal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni, Jonathan Herzig, Oran Lang, Eran Ofek, and Idan Szpektor · 2023
Earlier work this paper cites.
Hallucination of multimodal large language models: A survey
Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, and Mike Zheng Shou · 2024
Earlier work this paper cites.
Maplm: A real-world large-scale vision-language benchmark for map and traffic scene understanding
Xu Cao, Tong Zhou, Yunsheng Ma, Wenqian Ye, Can Cui, Kun Tang, Zhipeng Cao, Kaizhao Liang, Ziran Wang, James M Rehg, et al · 2024
Earlier work this paper cites.
Sam4mllm: Enhance multi-modal large language model for referring expression segmentation
Yi-Chia Chen, Wei-Hua Li, Cheng Sun, Yu-Chiang Frank Wang, and Chu-Song Chen · 2024
Earlier work this paper cites.
A survey on multimodal large language models for autonomous driving
Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, Yang Zhou, Kaizhao Liang, Jintai Chen, Juanwu Lu, Zichong Yang, Kuei-Da Liao, et al · 2024
Earlier work this paper cites.
Mapeval: A map-based evaluation of geo-spatial reasoning in foundation models
Mahir Labib Dihan, Md Tanvir Hassan, Md Tanvir Parvez, Md Hasebul Hasan, Md Almash Alam, Muhammad Aamir Cheema, Mohammed Eunus Ali, and Md Rizwan Parvez · 2024
Earlier work this paper cites.
Junqi Ge, Ziyi Chen, Jintao Lin, Jinguo Zhu, Xihui Liu, Jifeng Dai, and Xizhou Zhu · 2024
Earlier work this paper cites.
Rotary position embedding for vision transformer
Byeongho Heo, Song Park, Dongyoon Han, and Sangdoo Yun · 2024
Earlier work this paper cites.
Understanding the role of llms in multimodal evaluation benchmarks
Botian Jiang, Lei Li, Xiaonan Li, Zhaowei Li, Xiachong Feng, Lingpeng Kong, Qi Liu, and Xipeng Qiu · 2024
Earlier work this paper cites.
Lisa: Reasoning segmentation via large language model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia · 2024
Cited alongside, same era.
Ez-hoi: Vlm adaptation via guided prompt learning for zero-shot hoi detection
Qinqian Lei, Bo Wang, and Robby T. Tan · 2024
Cited alongside, same era.
Ocrbench: on the hidden mystery of ocr in large multimodal models
Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Cheng-Lin Liu, Lianwen Jin, and Xiang Bai · 2024
Cited alongside, same era.
Qvq: To see the world with wisdom
Qwen Team · 2024
Cited alongside, same era.
Pixellm: Pixel reasoning with large multimodal model
Zhongwei Ren, Zhicheng Huang, Yunchao Wei, Yao Zhao, Dongmei Fu, Jiashi Feng, and Xiaojie Jin · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Mergemix: A unified augmentation paradigm for visual and multi-modal understanding
Xin Jin, Siyuan Li, Siyong Jian, Kai Yu, and Huan Wang · 2025
Closest in time.
3d and 4d world modeling: A survey
Lingdong Kong, Wesley Yang, Jianbiao Mei, Youquan Liu, Ao Liang, Dekai Zhu, Dongyue Lu, Wei Yin, Xiaotao Hu, Mingkai Jia, et al · 2025
Closest in time.
Open R1 Multimodal
EvolvingLMMs Lab · 2025
Closest in time.
Hola: Zero-shot hoi detection with low-rank decomposed vlm feature adaptation
Qinqian Lei, Bo Wang, and Robby T. Tan · 2025
Closest in time.
Visual-rft: Visual reinforcement fine-tuning
Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, and Jiaqi Wang · 2025
Closest in time.
OpenAI o3 and o4-mini System Card
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al · 2024
Cited alongside, same era.
An empirical analysis on spatial reasoning capabilities of large multimodal models
Fatemeh Shiri, Xiao-Yu Guo, Mona Golestan Far, Xin Yu, Gholamreza Haffari, and Yuan-Fang Li · 2024
Cited alongside, same era.
V*: Guided visual search as a core mechanism in multimodal llms
Penghao Wu and Saining Xie · 2024
Cited alongside, same era.
Georeasoner: Reasoning on geospatially grounded context for natural language understanding
Yibo Yan and Joey Lee · 2024
Cited alongside, same era.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2024
Cited alongside, same era.
Qingbin Zeng, Qinglong Yang, Shunan Dong, Heming Du, Liang Zheng, Fengli Xu, and Yong Li · 2024
Cited alongside, same era.
Planagent: A multi-modal large language agent for closed-loop vehicle motion planning
Yupeng Zheng, Zebin Xing, Qichao Zhang, Bu Jin, Pengfei Li, Yuhang Zheng, Zhongpu Xia, Kun Zhan, Xianpeng Lang, Yaran Chen, et al · 2024
Cited alongside, same era.
OpenAI · 2025
Closest in time.
Skywork r1v: pioneering multimodal reasoning with chain-of-thought
Yi Peng, Xiaokun Wang, Yichen Wei, Jiangbo Pei, Weijie Qiu, Ai Jian, Yunzhuo Hao, Jiachun Pan, Tianyidan Xie, Li Ge, et al · 2025
Closest in time.
Vgrp-bench: Visual grid reasoning puzzle benchmark for large vision-language models
Yufan Ren, Konstantinos Tertikas, Shalini Maiti, Junlin Han, Tong Zhang, Sabine Süsstrunk, and Filippos Kokkinos · 2025
Closest in time.
Vlm-r1: A stable and generalizable r1-style large vision-language model
Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, et al · 2025
Closest in time.
Visualpuzzles: Decoupling multimodal reasoning evaluation from domain knowledge
Yueqi Song, Tianyue Ou, Yibo Kong, Zecheng Li, Graham Neubig, and Xiang Yue · 2025
Closest in time.
Reason-rft: Reinforcement fine-tuning for visual reasoning
Huajie Tan, Yuheng Ji, Xiaoshuai Hao, Minglan Lin, Pengwei Wang, Zhongyuan Wang, and Shanghang Zhang · 2025
Closest in time.
Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, et al · 2025
Closest in time.
Skywork r1v2: Multimodal hybrid reinforcement learning for reasoning
Yichen Wei, Yi Peng, Xiaokun Wang, Weijie Qiu, Wei Shen, Tianyidan Xie, Jiangbo Pei, Jianhao Zhang, Yunzhuo Hao, Xuchen Song, et al · 2025
Closest in time.
Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives
Shaoyuan Xie, Lingdong Kong, Yuhao Dong, Chonghao Sima, Wenwei Zhang, Qi Alfred Chen, Ziwei Liu, and Liang Pan · 2025
Closest in time.
Can large vision language models read maps like a human?
Shuo Xing, Zezhou Sun, Shuangyu Xie, Kaiyuan Chen, Yanjia Huang, Yuping Wang, Jiachen Li, Dezhen Song, and Zhengzhong Tu · 2025
Closest in time.
Geochain: Multimodal chain-of-thought for geographic reasoning
Sahiti Yerramilli, Nilay Pande, Rynaa Grover, and Jayant Sravan Tamarapalli · 2025
Closest in time.
Pyvision: Agentic vision with dynamic tooling
Shitian Zhao, Haoquan Zhang, Shaoheng Lin, Ming Li, Qilong Wu, Kaipeng Zhang, and Chen Wei · 2025
Closest in time.
dvoting: Fast voting for dllms
Sicheng Feng, Zigeng Chen, Xinyin Ma, Gongfan Fang, and Xinchao Wang · 2026
Closest in time.
Sponge tool attack: Stealthy denial-of-efficiency against tool-augmented agentic reasoning
Qi Li and Xinchao Wang · 2026
Closest in time.
Invisible safety threat: Malicious finetuning for llm via steganography
Guangnian Wan, Xinyin Ma, Gongfan Fang, and Xinchao Wang · 2026
Closest in time.
Logical phase transitions: Understanding collapse in llm logical reasoning
Xinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu, Wei Yang, and Zikai Song · 2026
Closest in time.