Fetching the paper…
Reading the bibliography…
Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment.
A formal basis for the heuristic determination of minimum cost paths
Peter E Hart, Nils J Nilsson, and Bertram Raphael. 1968 · 1968
Earlier work this paper cites.
Probing contextual language models for common ground with visual representations
Gabriel Ilharco, Rowan Zellers, Ali Farhadi, and Hannaneh Hajishirzi. 2021 · 2005
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020 · 2010
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2016 · 2016
Earlier work this paper cites.
Learning language games through interaction
Sida I. Wang, Percy Liang, and Christopher D. Manning. 2016 · 2016
Earlier work this paper cites.
ShapeWorld - a new test methodology for multimodal language understanding
Alexander Kuhnle and Ann Copestake. 2017 · 2017
Earlier work this paper cites.
Can language models encode perceptual structure without grounding? a case study in color
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard. 2021 · 2021
Earlier work this paper cites.
SpartQA: : A textual question answering benchmark for spatial reasoning
Roshanak Mirzaee, Hossein Rajaby Faghihi, Qiang Ning, and Parisa Kordjmashidi. 2021 · 2021
Earlier work this paper cites.
Mapping language models to grounded conceptual spaces
Roma Patel and Ellie Pavlick. 2021 · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022 · 2022
Earlier work this paper cites.
Look before you leap: Unveiling the power of GPT-4v in robotic vision-language planning
Yingdong Hu, Fanqi Lin, Tong Zhang, Li Yi, and Yang Gao. 2023 · 2023
Earlier work this paper cites.
What’s "up" with vision-language models? investigating their struggle with spatial reasoning
Amita Kamath, Jack Hessel, and Kai-Wei Chang. 2023a · 2023
Earlier work this paper cites.
Super-CLEVR: A virtual benchmark to diagnose domain robustness in visual reasoning
Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski, Wufei Ma, Benjamin Van Durme, and Alan Yuille. 2023 · 2023
Earlier work this paper cites.
Fangyu Liu, Guy Emerson, and Nigel Collier. 2023 · 2023
Earlier work this paper cites.
Linearly mapping from image to text space
Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick. 2023 · 2023
Earlier work this paper cites.
3d-aware visual question answering about parts, poses and occlusions
Xingrui Wang, Wufei Ma, Zhuowan Li, Adam Kortylewski, and Alan Yuille. 2023 · 2023
Earlier work this paper cites.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, and 1 others. 2023 · 2023
Cited alongside, same era.
Large language models for robotics: A survey
Fanlong Zeng, Wensheng Gan, Yongheng Wang, Ning Liu, and Philip S Yu. 2023 · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023 · 2023
Cited alongside, same era.
Mohamed Aghzal, Erion Plaku, and Ziyu Yao. 2024 · 2024
Cited alongside, same era.
Matteo G. Mecattaf, Ben Slater, Marko Tešić, Jonathan Prunty, Konstantinos Voudouris, and Lucy G. Cheke. 2024 · 2024
Later among the works it cites.
Sliding puzzles gym: A scalable benchmark for state representation in visual reinforcement learning
Bryan Lincoln Marques de Oliveira, Bruno Brandão, Murilo Lopes da Luz, Luana Guedes Barros Martins, Telma Woerle de Lima Soares, and Luckeciano Carvalho Melo. 2024 · 2024
Later among the works it cites.
OpenAI, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander Mądry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, Alex Nichol, and 400 others. 2024 · 2024
Later among the works it cites.
Md Imbesat Hassan Rizvi, Xiaodan Zhu, and Iryna Gurevych. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bernd Bohnet, Azade Nova, Aaron T. Parisi, Kevin Swersky, Katayoon Goshvadi, Hanjun Dai, Dale Schuurmans, Noah Fiedel, and Hanie Sedghi. 2024 · 2024
Cited alongside, same era.
An introduction to vision-language modeling
Florian Bordes, Richard Yuanzhe Pang, Anurag Ajay, Alexander C. Li, Adrien Bardes, Suzanne Petryk, Oscar Mañas, Zhiqiu Lin, Anas Mahmoud, Bargav Jayaraman, Mark Ibrahim, Melissa Hall, Yunyang Xiong, Jonathan Lebensold, Candace Ross, Srihari Jayakumar, Chuan Guo, Diane Bouchacourt, Haider Al-Tahan, and 22 others. 2024 · 2024
Cited alongside, same era.
Understanding the limits of vision language models through the lens of the binding problem
Declan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata, Kia Ghods, Amogh Joshi, Alexander Ku, Steven M. Frankland, Thomas L. Griffiths, Jonathan D. Cohen, and Taylor W. Webb. 2024 · 2024
Cited alongside, same era.
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, and 1 others. 2024 · 2024
Cited alongside, same era.
SpatialRGPT: Grounded spatial reasoning in vision language models
An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu. 2024 · 2024
Cited alongside, same era.
Introducing the next generation of claude
Claude Team. 2024 · 2024
Cited alongside, same era.
PUZZLES: A benchmark for neural algorithmic reasoning
Benjamin Estermann, Luca A. Lanzendörfer, Yannick Niedermayr, and Roger Wattenhofer. 2024 · 2024
Cited alongside, same era.
Gemini 2.0 flash (experimental)
Gemini Team. 2024 · 2024
Cited alongside, same era.
Later among the works it cites.
Smart vision-language reasoners
Denisa Roberts and Lucas Roberts. 2024 · 2024
Later among the works it cites.
Ying Su, Zhan Ling, Haochen Shi, Jiayang Cheng, Yauwai Yim, and Yangqiu Song. 2024 · 2024
Later among the works it cites.
Yihong Tang, Ao Qu, Zhaokai Wang, Dingyi Zhuang, Zhaofeng Wu, Wei Ma, Shenhao Wang, Yunhan Zheng, Zhan Zhao, and Jinhua Zhao. 2024 · 2024
Later among the works it cites.
VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs
Qiucheng Wu, Handong Zhao, Michael Saxon, Trung Bui, William Yang Wang, Yang Zhang, and Shiyu Chang. 2024 · 2024
Later among the works it cites.
Evaluating spatial understanding of large language models
Yutaro Yamada, Yihan Bao, Andrew K. Lampinen, Jungo Kasai, and Ilker Yildirim. 2024 · 2024
Later among the works it cites.
Wenqi Zhang, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu, and Yueting Zhuang. 2024 · 2024
Later among the works it cites.
Understanding the limits of vision language models through the lens of the binding problem
Declan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata, Kia Ghods, Amogh Joshi, Alexander Ku, Steven M. Frankland, Thomas L. Griffiths, Jonathan D. Cohen, and Taylor W. Webb. 2025 · 2025
Closest in time.
Lin Duan, Yanming Xiu, and Maria Gorlatova. 2025 · 2025
Closest in time.
Reflective planning: Vision-language models for multi-stage long-horizon robotic manipulation
Yunhai Feng, Jiaming Han, Zhuoran Yang, Xiangyu Yue, Sergey Levine, and Jianlan Luo. 2025 · 2025
Closest in time.
Are large vision language models good game players?
Xinyu Wang, Bohan Zhuang, and Qi Wu. 2025 · 2025
Closest in time.
Baining Zhao, Ziyou Wang, Jianjie Fang, Chen Gao, Fanhang Man, Jinqiang Cui, Xin Wang, Xinlei Chen, Yong Li, and Wenwu Zhu. 2025 · 2025
Closest in time.