Fetching the paper…
Reading the bibliography…
Spatial relation reasoning is a crucial task for multimodal large language models (MLLMs) to understand the objective world.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Where am i? who am i? the relation between spatial cognition, social cognition and individual differences in the built environment
Michael J Proulx, Orlin S Todorov, Amanda Taylor Aiken, and Alexandra A de Sousa. 2016 · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Earlier work this paper cites.
A corpus of natural language for visual reasoning
Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi. 2017 · 2017
Earlier work this paper cites.
Spatialvoc2k: A multilingual dataset of images with annotations and features for spatial relations between objects
Anja Belz, Adrian Muscat, Pierre Anguill, Mouhamadou Sow, Gaétan Vincent, and Yassine Zinessabah. 2018 · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. 2019 · 2019
Earlier work this paper cites.
Rel3d: A minimally contrastive benchmark for grounding spatial relations in 3d
Ankit Goyal, Kaiyu Yang, Dawei Yang, and Jia Deng. 2020 · 2020
Earlier work this paper cites.
What explains the relationship between spatial and mathematical skills? a review of evidence from brain and behavior
Zachary Hawes and Daniel Ansari. 2020 · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2021 · 2021
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Can vision-language models think from a first-person perspective?
Sijie Cheng, Zhicheng Guo, Jingwen Wu, Kechen Fang, Peng Li, Huaping Liu, and Yang Liu. 2023 · 2023
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, and Rongrong Ji. 2023 · 2023
Cited alongside, same era.
Texts as images in prompt tuning for multi-label image recognition
Zixian Guo, Bowen Dong, Zhilong Ji, Jinfeng Bai, Yiwen Guo, and Wangmeng Zuo. 2023 · 2023
Cited alongside, same era.
Spatialbot: Precise spatial understanding with vision language models
Wenxiao Cai, Yaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li, Wankou Yang, Hao Dong, and Bo Zhao. 2024 · 2024
Later among the works it cites.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. 2024 · 2024
Later among the works it cites.
Spatialrgpt: Grounded spatial reasoning in vision language model
An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu. 2024 · 2024
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hugo Laurençon, Daniel van Strien, Stas Bekman, Leo Tronchon, Lucile Saulnier, Thomas Wang, Siddharth Karamcheti, Amanpreet Singh, Giada Pistilli, Yacine Jernite, et al. 2023 · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Cited alongside, same era.
Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning
Mustafa Shukor, Alexandre Rame, Corentin Dancette, and Matthieu Cord. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023 · 2023
Cited alongside, same era.
Representation-enhanced status replay network for multisource remote-sensing image classification
Junjie Wang, Wei Li, Yinjian Wang, Ran Tao, and Qian Du. 2023 · 2023
Cited alongside, same era.
mplug-owl: Modularization empowers large language models with multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. 2023 · 2023
Cited alongside, same era.
Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models
Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang. 2023 · 2023
Cited alongside, same era.
Visual spatial reasoning
Fangyu Liu, Guy Emerson, and Nigel Collier. 2023a
Cited in the paper.
Mengfei Du, Binhao Wu, Zejun Li, Xuanjing Huang, and Zhongyu Wei. 2024 · 2024
Later among the works it cites.
A survey for foundation models in autonomous driving
Haoxiang Gao, Yaqian Li, Kaiwen Long, Ming Yang, and Yiqing Shen. 2024 · 2024
Later among the works it cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024 · 2024
Later among the works it cites.
T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering
Lei Wang, Yi Hu, Jiabang He, Xing Xu, Ning Liu, Hui Liu, and Heng Tao Shen. 2024 · 2024
Later among the works it cites.
Can transformers capture spatial relations between objects?
Chuan Wen, Dinesh Jayaraman, and Yang Gao. 2024 · 2024
Later among the works it cites.
Cocot: Contrastive chain-of-thought prompting for large multimodal models with multiple image inputs
Daoan Zhang, Junming Yang, Hanjia Lyu, Zijian Jin, Yuan Yao, Mingkai Chen, and Jiebo Luo. 2024 · 2024
Later among the works it cites.
Spatialsense: An adversarially crowdsourced benchmark for spatial relation recognition
Kaiyu Yang, Olga Russakovsky, and Jia Deng. 2019 · 2060
Closest in time.