Fetching the paper…
Reading the bibliography…
The ability to understand and reason about spatial relationships between objects in images is an important component of visual reasoning.
Lxmert: Learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M. (2019) · 1908
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., · 1910
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S., Ordonez, V., Matten, M., and Berg, T. (2014) · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. (2015) · 2015
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., · 2017
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D. (2019) · 2019
Earlier work this paper cites.
Spatialsense: An adversarially crowdsourced benchmark for spatial relation recognition
Yang, K., Russakovsky, O., and Deng, J. (2019) · 2019
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., · 2021
Earlier work this paper cites.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Zeng, Y., Zhang, X., and Li, H. (2021) · 2021
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S. (2022) · 2022
Earlier work this paper cites.
Liu, F., Emerson, G., and Collier, N. (2022) · 2022
Cited alongside, same era.
Flava: A foundational language and vision alignment model
Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., and Kiela, D. (2022) · 2022
Cited alongside, same era.
Coca: Contrastive captioners are image-text foundation models
Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., and Wu, Y. (2022) · 2022
Cited alongside, same era.
When and why vision-language models behave like bag-of-words models, and what to do about it?
Yuksekgonul, M., Bianchi, F., Kalluri, P., Jurafsky, D., and Zou, J. (2022) · 2022
Cited alongside, same era.
Zoedepth: Zero-shot transfer by combining relative and metric depth
Towards grounded visual spatial reasoning in multi-modal vision language models
Rajabi, N. and Kosecka, J. (2023) · 2023
Later among the works it cites.
Large language models are not fair evaluators
Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., and Sui, Z. (2023) · 2023
Later among the works it cites.
Large language models are not robust multiple choice selectors
Zheng, C., Zhou, H., Meng, F., Zhou, J., and Huang, M. (2023) · 2023
Later among the works it cites.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Chen, Z., Wang, W., Tian, H., Ye, S., Gao, Z., Cui, E., Tong, W., Hu, K., Luo, J., Ma, Z., · 2024
Closest in time.
Llava-next: Stronger llms supercharge multimodal capabilities in the wild
Li, B., Zhang, K., Zhang, H., Guo, D., Zhang, R., Li, F., Zhang, Y., Liu, Z., and Li, C. (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bhat, S. F., Birkl, R., Wofk, D., Wonka, P., and Müller, M. (2023) · 2023
Cited alongside, same era.
What’s “up” with vision-language models? investigating their struggle with spatial reasoning
Kamath, A., Hessel, J., and Chang, K.-W. (2023) · 2023
Cited alongside, same era.
Segment anything
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., · 2023
Cited alongside, same era.
Li, J., Li, D., Savarese, S., and Hoi, S. (2023) · 2023
Cited alongside, same era.
Large language models sensitivity to the order of options in multiple-choice questions
Pezeshkpour, P. and Hruschka, E. (2023) · 2023
Cited alongside, same era.
Visual spatial reasoning
Liu, F., Emerson, G., and Collier, N. (2023a)
Cited in the paper.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J. (2023b)
Cited in the paper.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J.,
Cited in the paper.
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J. (2024) · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., · 2024
Closest in time.
Wang, X., Hu, C., Ma, B., Röttger, P., and Plank, B. (2024) · 2024
Closest in time.
Strengthened symbol binding makes large language models reliable multiple-choice selectors
Xue, M., Hu, Z., Zhao, M., Liu, L., Liao, K., Li, S., Han, H., and Yin, C. (2024) · 2024
Closest in time.