Fetching the paper…
Reading the bibliography…
Vision-Language Models (VLMs) have recently emerged as powerful tools, excelling in tasks that integrate visual and textual comprehension, such as image captioning, visual question answering, and image-text retrieval.
Mental rotation of random two-dimensional shapes
Lynn A Cooper · 1975
Earlier work this paper cites.
Mental rotations of the alphabet letters
Steven G Vandenberg and Allan R Kuse · 1975
Earlier work this paper cites.
Manual for kit of factor-referenced cognitive tests, 1976
Ruth B Ekstrom and Harry Horace Harman · 1976
Earlier work this paper cites.
Human spatial abilities: psychometric studies and environmental, genetic, hormonal, and neurological influences
Meredith G McGee · 1979
Earlier work this paper cites.
Emergence and characterization of sex differences in spatial ability: A meta-analysis
Marcia C Linn and Anne C Petersen · 1985
Earlier work this paper cites.
The Body in the Mind: The Bodily Basis of Meaning, Imagination, and Reason
Mark Johnson · 1987
Earlier work this paper cites.
Human cognitive abilities: A survey of factor-analytic studies
John B Carroll · 1993
Earlier work this paper cites.
Scale and multiple psychologies of space
Daniel R Montello · 1993
Earlier work this paper cites.
Spatial working memory, visual attention, and object identity in visual search
Gordon D Logan · 1996
Earlier work this paper cites.
Qualitative spatial reasoning: ontologies, granularity and containment
Brandon Bennett · 1998
Earlier work this paper cites.
Spatial Cognition: An Interdisciplinary Approach to Representing and Processing Spatial Knowledge
Roberta L Klatzky · 1998
Earlier work this paper cites.
Cognitive correlates of spatial ability: Evidence for two distinct processes
James W Pellegrino, Robert Kail, and Brian McCloskey · 1998
Earlier work this paper cites.
Spatial orientation and wayfinding in large-scale virtual spaces
Rudolph P Darken, Terry Allard, and Lisa B Achille · 1999
Earlier work this paper cites.
Mental rotation of photographs of human bodies
Meredith Wraga, Roger N Shepard, Patricia S Churchland, Souheil Inati, and Stephen M Kosslyn · 2000
Earlier work this paper cites.
Human spatial representation: insights from animals
Ranxiao Wang and Elizabeth Spelke · 2002
Earlier work this paper cites.
Greedy decoding for statistical machine translation in almost linear time
Ulrich Germann · 2003
Earlier work this paper cites.
3d user interfaces: Theory and practice
Doug A Bowman, Ernst Kruijff, Joseph J LaViola Jr, and Ivan Poupyrev · 2004
Earlier work this paper cites.
A dissociation between mental rotation and perspective-taking spatial abilities
Mary Hegarty and David Waller · 2004
Earlier work this paper cites.
Cognitive components of environmental spatial cognition
Daniel R Montello · 2005
Earlier work this paper cites.
Probabilistic robotics
Sebastian Thrun, Wolfram Burgard, and Dieter Fox · 2005
Cited alongside, same era.
Spatial reference in linguistic human-robot interaction: Iterative, empirically supported development of a model of projective relations
Reinhard Moratz and Thora Tenbrink · 2006
Cited alongside, same era.
Space and the parietal cortex
Masud Husain and Parashkev Nachev · 2007
Cited alongside, same era.
Spatial cognition and the brain
Neil Burgess · 2008
Cited alongside, same era.
Spatial and temporal reasoning
Oliviero Stock · 2008
Cited alongside, same era.
What does the mental rotation test measure? an analysis of item difficulty and item characteristics
Andre F Caissie, Francois Vigneau, and Douglas A Bors · 2009
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning, 2023
Tri Dao · 2023
Later among the works it cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Later among the works it cites.
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models, 2024
An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spatial ability for stem domains: aligning over 50 years of cumulative psychological knowledge solidifies its importance
Jonathan Wai, David Lubinski, and Camilla P Benbow · 2009
Cited alongside, same era.
Chapter 7 - components of spatial intelligence
Mary Hegarty · 2010
Cited alongside, same era.
Spatial relations among parts of objects: a survey
Christian Freksa, Alexander Klippel, Louisa Knuf, Bernhard Nebel, and Stefan Wölfl · 2013
Cited alongside, same era.
The malleability of spatial skills: a meta-analysis of training studies
David H Uttal, Nathaniel G Meadow, Elizabeth Tipton, Linda L Hand, Alison R Alden, Christopher Warren, and Nora S Newcombe · 2013
Cited alongside, same era.
Spatial training improves children’s mathematics ability
Yi-Ling Cheng and Kelly S Mix · 2014
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Later among the works it cites.
Ji Hyeok Jung, Eun Tae Kim, Seo Yeon Kim, Joo Ho Lee, Bumsoo Kim, and Buru Chang · 2024
Later among the works it cites.
Building and better understanding vision-language models: insights and future directions, 2024
Hugo Laurençon, Andrés Marafioti, Victor Sanh, and Léo Tronchon · 2024
Later among the works it cites.
Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models
Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li · 2024
Later among the works it cites.
Smolvlm - small yet mighty vision language model, 2024
Andres Marafioti, Merve Noyan, Miquel Farré, Elie Bakouch, and Pedro Cuenca · 2024
Later among the works it cites.
Towards grounded visual spatial reasoning in multi-modal vision language models
Navid Rajabi and Jana Kosecka · 2024
Later among the works it cites.
An empirical analysis on spatial reasoning capabilities of large multimodal models
Fatemeh Shiri, Xiao-Yu Guo, Mona Golestan Far, Xin Yu, Reza Haf, and Yuan-Fang Li · 2024
Later among the works it cites.
Yihong Tang, Ao Qu, Zhaokai Wang, Dingyi Zhuang, Zhaofeng Wu, Wei Ma, Shenhao Wang, Yunhan Zheng, Zhan Zhao, and Jinhua Zhao · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al · 2024
Later among the works it cites.
Introducing stable diffusion 3.5
Stability AI · 2025
Closest in time.
Spatialrgpt: Grounded spatial reasoning in vision-language models
An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu · 2025
Closest in time.
Is a picture worth a thousand words? delving into spatial reasoning for vision language models
Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, Sharon Li, and Neel Joshi · 2025
Closest in time.
Wenrui Xu, Dalin Lyu, Weihang Wang, Jie Feng, Chen Gao, and Yong Li · 2025
Closest in time.