Fetching the paper…
Reading the bibliography…
Spatial understanding is essential for Multimodal Large Language Models (MLLMs) to support perception, reasoning, and planning in embodied environments.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Kolve, E.; Mottaghi, R.; Han, W.; VanderBilt, E.; Weihs, L.; Herrasti, A.; Deitke, M.; Ehsani, K.; Gordon, D.; Zhu, Y.; et al. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Earlier work this paper cites.
What’s “up” with vision-language models? Investigating their struggle with spatial reasoning
Kamath, A.; Hessel, J.; and Chang, K.-W. 2023 · 2023
Earlier work this paper cites.
ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation
Li, X.; Zhang, M.; Geng, Y.; Geng, H.; Long, Y.; Shen, Y.; Zhang, R.; Liu, J.; and Dong, H. 2023 · 2023
Earlier work this paper cites.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Earlier work this paper cites.
Xu, L.; Xie, H.; Qin, S.-Z. J.; Tao, X.; and Wang, F. L. 2023 · 2023
Earlier work this paper cites.
Multi-Object Hallucination in Vision-Language Models
Chen, X.; Ma, Z.; Zhang, X.; Xu, S.; Qian, S.; Yang, J.; Fouhey, D. F.; and Chai, J. 2024 · 2024
Earlier work this paper cites.
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
Du, M.; Wu, B.; Li, Z.; Huang, X.; and Wei, Z. 2024 · 2024
Cited alongside, same era.
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
Fei, H.; Wu, S.; Ji, W.; Zhang, H.; Zhang, M.; Lee, M.-L.; and Hsu, W. 2024 · 2024
Cited alongside, same era.
Rotary position embedding for vision transformer
Heo, B.; Park, S.; Han, D.; and Yun, S. 2024 · 2024
Cited alongside, same era.
Hurst, A.; Lerer, A.; Goucher, A. P.; Perelman, A.; Ramesh, A.; Clark, A.; Ostrow, A.; Welihinda, A.; Hayes, A.; Radford, A.; et al. 2024 · 2024
Cited alongside, same era.
Llava-onevision: Easy visual task transfer
Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Li, Y.; Liu, Z.; and Li, C. 2024 · 2024
Cited alongside, same era.
Video-R1: Reinforcing Video Reasoning in MLLMs
Feng, K.; Gong, K.; Li, B.; Guo, Z.; Wang, Y.; Peng, T.; Wu, J.; Zhang, X.; Wang, B.; and Yue, X. 2025 · 2025
Closest in time.
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
Jia, M.; Qi, Z.; Zhang, S.; Zhang, W.; Yu, X.; He, J.; Wang, H.; and Yi, L. 2025 · 2025
Closest in time.
”Well, Keep Thinking”: Enhancing LLM Reasoning with Adaptive Injection Decoding
Jin, H.; Yeom, J. W.; Bae, S.; and Kim, T. 2025 · 2025
Closest in time.
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
Li, C.; Wu, W.; Zhang, H.; Xia, Y.; Mao, S.; Dong, L.; Vulić, I.; and Wei, F. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
World model on million-length video and language with ringattention
Liu, H.; Yan, W.; Zaharia, M.; and Abbeel, P. 2024 · 2024
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J.; Ahmed, M.; Lu, Y.; Pan, S.; Bo, W.; and Liu, Y. 2024 · 2024
Cited alongside, same era.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Tong, S.; Liu, Z.; Zhai, Y.; Ma, Y.; LeCun, Y.; and Xie, S. 2024 · 2024
Cited alongside, same era.
AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
Xiong, C.; Shen, C.; Li, X.; Zhou, K.; Liu, J.; Wang, R.; and Dong, H. 2024 · 2024
Cited alongside, same era.
An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models
Zhou, X.; He, J.; Ke, Y.; Zhu, G.; Gutierrez Basulto, V.; and Pan, J. 2024 · 2024
Cited alongside, same era.
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; Zhong, H.; Zhu, Y.; Yang, M.; Li, Z.; Wan, J.; Wang, P.; Ding, W.; Fu, Z.; Xu, Y.; Ye, J.; Zhang, X.; Xie, T.; Cheng, Z.; Zhang, H.; Yang, Z.; Xu, H.; and Lin, J. 2025 · 2025
Cited alongside, same era.
Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
Chen, S.; Zhu, T.; Zhou, R.; Zhang, J.; Gao, S.; Niebles, J. C.; Geva, M.; He, J.; Wu, J.; and Li, M. 2025 · 2025
Cited alongside, same era.
Liao, Z.; Xie, Q.; Zhang, Y.; Kong, Z.; Lu, H.; Yang, Z.; and Deng, Z. 2025 · 2025
Closest in time.
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
Lin, T.; Li, G.; Zhong, Y.; Zou, Y.; and Zhao, B. 2025 · 2025
Closest in time.
Liu, Y.; Chi, D.; Wu, S.; Zhang, Z.; Hu, Y.; Zhang, L.; Zhang, Y.; Wu, S.; Cao, T.; Huang, G.; Huang, H.; Tian, G.; Qiu, W.; Quan, X.; Hao, J.; and Zhuang, Y. 2025 · 2025
Closest in time.
Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training
Luo, G.; Yang, X.; Dou, W.; Wang, Z.; Liu, J.; Dai, J.; Qiao, Y.; and Zhu, X. 2025 · 2025
Closest in time.
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
Ouyang, K.; Liu, Y.; Wu, H.; Liu, Y.; Zhou, H.; Zhou, J.; Meng, F.; and Sun, X. 2025 · 2025
Closest in time.
VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Wei, X.; Liu, X.; Zang, Y.; Dong, X.; Zhang, P.; Cao, Y.; Tong, J.; Duan, H.; Guo, Q.; Wang, J.; Qiu, X.; and Lin, D. 2025 · 2025
Closest in time.
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Yang, J.; Yang, S.; Gupta, A. W.; Han, R.; Fei-Fei, L.; and Xie, S. 2025 · 2025
Closest in time.
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
Yeh, C.-H.; Wang, C.; Tong, S.; Cheng, T.-Y.; Wang, R.; Chu, T.; Zhai, Y.; Chen, Y.; Gao, S.; and Ma, Y. 2025 · 2025
Closest in time.
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
Zhang, J.; Khayatkhoei, M.; Chhikara, P.; and Ilievski, F. 2025 · 2025
Closest in time.
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Zheng, D.; Huang, S.; and Wang, L. 2025 · 2025
Closest in time.
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Zhu, J.; Wang, W.; Chen, Z.; Liu, Z.; Ye, S.; Gu, L.; Tian, H.; Duan, Y.; Su, W.; Shao, J.; Gao, Z.; Cui, E.; Wang, X.; Cao, Y.; Liu, Y.; Wei, X.; Zhang, H.; Wang, H.; Xu, W.; Li, H.; Wang, J.; Deng, N.; Li, S.; He, Y.; Jiang, T.; Luo, J.; Wang, Y.; He, C.; Shi, B.; Zhang, X.; Shao, W.; He, J.; Xiong, Y.; Qu, W.; Sun, P.; Jiao, P.; Lv, H.; Wu, L.; Zhang, K.; Deng, H.; Ge, J.; Chen, K.; Wang, L.; Dou, M.; Lu, L.; Zhu, X.; Lu, T.; Lin, D.; Qiao, Y.; Dai, J.; and Wang, W. 2025 · 2025
Closest in time.