Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in general vision-language tasks.
Cognitive maps in rats and men
Edward~C Tolman · 1948
Earlier work this paper cites.
Human spatial abilities: psychometric studies and environmental, genetic, hormonal, and neurological influences
Mark~G McGee · 1979
Earlier work this paper cites.
Spatial imagery in deductive reasoning: a functional mri study
Markus Knauff, Thomas Mulack, Jan Kassubek, Helmut~R Salih, and Mark~W Greenlee · 2002
Earlier work this paper cites.
Localizing age-related individual differences in a hierarchical structure
Timothy~A Salthouse · 2004
Earlier work this paper cites.
Spatial and object visualization cognitive styles: Validation studies in 3800 individuals
Christopher~F Chabris, Thomas~E Jerde, Anita~W Woolley, Margaret~E Gerbasi, Jonathon~P Schuldt, Sean~L Bennett, J~Richard Hackman, and Stephen~M Kosslyn · 2006
Earlier work this paper cites.
Place cells, grid cells, and the brain's spatial representation system
Edvard~I Moser, Emilio Kropff, and May-Britt Moser · 2008
Earlier work this paper cites.
On the development and measurement of spatial ability
H~Bayram Yılmaz · 2009
Earlier work this paper cites.
Spatial representation across species: Geometry, language, and maps
Barbara Landau and Laura Lakusta · 2009
Earlier work this paper cites.
The neuroscience of human intelligence differences
Ian~J. Deary, Lars Penke, and Wendy Johnson · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara Berg · 2011
Earlier work this paper cites.
Handbook of spatial cognition
David~Ed Waller and Lynn~Ed Nadel · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C~Lawrence Zitnick · 2014
Earlier work this paper cites.
Zipf’s word frequency law in natural language: A critical review and future directions
Steven~T. Piantadosi · 2014
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2020
Earlier work this paper cites.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Yan Zeng, Xinsong Zhang, and Hang Li · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze~Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Earlier work this paper cites.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira · 2021
Earlier work this paper cites.
Spartqa:: A textual question answering benchmark for spatial reasoning
Roshanak Mirzaee, Hossein~Rajaby Faghihi, Qiang Ning, and Parisa Kordjmashidi · 2021
Earlier work this paper cites.
Learning to generate scene graph from natural language supervision
Yiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu, and Yin Li · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et~al · 2022
Earlier work this paper cites.
Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts
Zhengxiang Shi, Qiang Zhang, and Aldo Lipani · 2022
Earlier work this paper cites.
Transfer learning with synthetic corpora for spatial role labeling and reasoning
Roshanak Mirzaee and Parisa Kordjamshidi · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et~al · 2022
Earlier work this paper cites.
Your transformer may not be as powerful as you expect
Shengjie Luo, Shanda Li, Shuxin Zheng, Tie-Yan Liu, Liwei Wang, and Di~He · 2022
Earlier work this paper cites.
Petr: Position embedding transformation for multi-view 3d object detection
Yingfei Liu, Tiancai Wang, Xiangyu Zhang, and Jian Sun · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed~H. Chi, Quoc~V Le, and Denny Zhou · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia~Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et~al · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos~E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu~Lin, Jinrui Yang, Xiawu Zheng, Ke~Li, Xing Sun, et~al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin~Tat Lee, Yuanzhi Li, Scott Lundberg, et~al · 2023
Cited alongside, same era.
Does spatial cognition emerge in frontier models?
Santhosh~Kumar Ramakrishnan, Erik Wijmans, Philipp Kraehenbuehl, and Vladlen Koltun · 2024
Later among the works it cites.
Lmdrive: Closed-loop end-to-end driving with large language models
Hao Shao, Yuxuan Hu, Letian Wang, Guanglu Song, Steven~L Waslander, Yu~Liu, and Hongsheng Li · 2024
Later among the works it cites.
Holistic autonomous driving understanding by bird's-eye-view injected multi-modal large models
Xinpeng Ding, Jianhua Han, Hang Xu, Xiaodan Liang, Wei Zhang, and Xiaomeng Li · 2024
Later among the works it cites.
To preserve or to compress: An in-depth study of connector selection in multimodal large language models
Junyan Lin, Haoran Chen, Dawei Zhu, and Xiaoyu Shen · 2024
Later among the works it cites.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong~Jae Lee · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visual spatial reasoning
Fangyu Liu, Guy Emerson, and Nigel Collier · 2023
Cited alongside, same era.
3d-llm: Injecting the 3d world into large language models
Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan · 2023
Cited alongside, same era.
A survey on multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke~Li, Xing Sun, Tong Xu, and Enhong Chen · 2023
Cited alongside, same era.
What's ``up'' with vision-language models? investigating their struggle with spatial reasoning
Amita Kamath, Jack Hessel, and Kai-Wei Chang · 2023
Cited alongside, same era.
A benchmark for reasoning with spatial prepositions
Iulia Comsa and Srini Narayanan · 2023
Cited alongside, same era.
Scene as occupancy
Wenwen Tong, Chonghao Sima, Tai Wang, Li~Chen, Silei Wu, Hanming Deng, Yi~Gu, Lewei Lu, Ping Luo, Dahua Lin, et~al · 2023
Cited alongside, same era.
Uni3d: Exploring unified 3d representation at scale
Junsheng Zhou, Jinsheng Wang, Baorui Ma, Yu-Shen Liu, Tiejun Huang, and Xinlong Wang · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Cited alongside, same era.
Honeybee: Locality-enhanced projector for multimodal llm
Junbum Cha, Wooyoung Kang, Jonghwan Mun, and Byungseok Roh · 2024
Later among the works it cites.
Nvlm: Open frontier-class multimodal llms
Wenliang Dai, Nayeon Lee, Boxin Wang, Zhuolin Yang, Zihan Liu, Jon Barker, Tuomas Rintamaki, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai~Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et~al · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu~Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Later among the works it cites.
Paligemma: A versatile 3b vlm for transfer
Lucas Beyer, Andreas Steiner, André~Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, et~al · 2024
Later among the works it cites.
Paligemma 2: A family of versatile vlms for transfer
Andreas Steiner, André~Susano Pinto, Michael Tschannen, Daniel Keysers, Xiao Wang, Yonatan Bitton, Alexey Gritsenko, Matthias Minderer, Anthony Sherbondy, Shangbang Long, et~al · 2024
Later among the works it cites.
Analyzing the language of visual tokens
David~M Chan, Rodolfo Corona, Joonyong Park, Cheol~Jun Cho, Yutong Bai, and Trevor Darrell · 2024
Later among the works it cites.
Neuro-symbolic training for reasoning over spatial language
Tanawan Premsri and Parisa Kordjamshidi · 2024
Later among the works it cites.
Enhancing image layout control with loss-guided diffusion models
Zakaria Patel and Kirill Serkh · 2024
Later among the works it cites.
Md~Imbesat~Hassan Rizvi, Xiaodan Zhu, and Iryna Gurevych · 2024
Later among the works it cites.
3dsrbench: A comprehensive 3d spatial reasoning benchmark
Wufei Ma, Haoyu Chen, Guofeng Zhang, Celso~M de~Melo, Jieneng Chen, and Alan Yuille · 2024
Later among the works it cites.
Robopoint: A vision-language model for spatial affordance prediction for robotics
Wentao Yuan, Jiafei Duan, Valts Blukis, Wilbert Pumacay, Ranjay Krishna, Adithyavairavan Murali, Arsalan Mousavian, and Dieter Fox · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
Mind's eye of LLMs: Visualization-of-thought elicits spatial reasoning in large language models
Wenshan Wu, Shaoguang Mao, Yadong Zhang, Yan Xia, Li~Dong, Lei Cui, and Furu Wei · 2024
Later among the works it cites.
Spatialbot: Precise spatial understanding with vision language models, 2024
Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li, Wankou Yang, Hao Dong, and Bo~Zhao · 2024
Later among the works it cites.
RoboSpatial: Teaching spatial understanding to 2D and 3D vision-language models for robotics
Chan~Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree, Yu~Su, and Stan Birchfield · 2025
Closest in time.
Fp3: A 3d foundation policy for robotic manipulation
Rujia Yang, Geng Chen, Chuan Wen, and Yang Gao · 2025
Closest in time.
ivispar–an interactive visual-spatial reasoning benchmark for vlms
Julius Mayer, Mohamad Ballout, Serwan Jassim, Farbod~Nosrat Nezami, and Elia Bruni · 2025
Closest in time.
Yuecheng Liu, Dafeng Chi, Shiguang Wu, Zhanguang Zhang, Yaochen Hu, Lingfeng Zhang, Yingxue Zhang, Shuang Wu, Tongtong Cao, Guowei Huang, et~al · 2025
Closest in time.
Vlm-r1: A stable and generalizable r1-style large vision-language model
Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, et~al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et~al · 2025
Closest in time.
MLLMs know where to look: Training-free perception of small visual details with multimodal LLMs
Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski · 2025
Closest in time.
Spatial thinking as the missing piece in mathematics curricula
Katie~A. Gilligan-Lee, Zachary C.~K. Hawes, and Kelly~S. Mix · 2056
Closest in time.