Fetching the paper…
Reading the bibliography…
Geometric shapes play important roles in both physical world and human cognition.
The shape of you: do individuals associate particular geometric shapes with identity?
Manippa, V. and Tommasi, L · 1936
Earlier work this paper cites.
An image synthesizer
Perlin, K · 1985
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Computational geometry: algorithms and applications, 3rd Edition
de Berg, M., Cheong, O., van Kreveld, M. J., and Overmars, M. H · 2008
Earlier work this paper cites.
From natural geometry to spatial cognition
Tommasi, L., Chiandetti, C., Pecchia, T., Sovrano, V. A., and Vallortigara, G · 2011
Earlier work this paper cites.
Efficient scale- and rotation-invariant encoding of visual words for image classification
Anwar, H., Zambanini, S., and Kampel, M · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., Salvato, M., Kolve, E., Seo, M. J., Hajishirzi, H., and Farhadi, A · 2016
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Agrawal, A., Summers-Stay, D., Batra, D., and Parikh, D · 2019
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Towards VQA models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Pathvqa: 30000+ questions for medical visual question answering
He, X., Zhang, Y., Mou, L., Xing, E., and Xie, P · 2020
Earlier work this paper cites.
Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning
Chen, J., Tang, J., Qin, J., Liang, X., Liu, L., Xing, E. P., and Lin, L · 2021
Earlier work this paper cites.
Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering
Liu, B., Zhan, L.-M., Xu, L., Ma, L., Yang, Y., and Wu, X.-M · 2021
Earlier work this paper cites.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Lu, P., Gong, R., Jiang, S., Qiu, L., Huang, S., Liang, X., and Zhu, S · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Mathew, M., Karatzas, D., and Jawahar, C · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression
Chen, J., Li, T., Qin, J., Lu, P., Lin, L., Chen, C., and Liang, X · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K., Zhu, S., Tafjord, O., Clark, P., and Kalyan, A · 2022
Earlier work this paper cites.
Winoground: Probing vision and language models for visio-linguistic compositionality, 2022
Thrush, T., Jiang, R., Bartolo, M., Singh, A., Williams, A., Kiela, D., and Ross, C · 2022
Earlier work this paper cites.
Reproducible scaling laws for contrastive language-image learning
Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J · 2023
Earlier work this paper cites.
Opencompass: A universal evaluation platform for foundation models
Contributors, O · 2023
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. C. H · 2023
Earlier work this paper cites.
MME: A comprehensive evaluation benchmark for multimodal large language models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Qiu, Z., Lin, W., Yang, J., Zheng, X., Li, K., Sun, X., and Ji, R · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Google · 2023
Cited alongside, same era.
Geomverse: A systematic evaluation of large models for geometric reasoning
Kazemi, M., Alvari, H., Anand, A., Wu, J., Chen, X., and Soricut, R · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J · 2023
Cited alongside, same era.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Cited alongside, same era.
Toward a holistic evaluation of robustness in clip models, 2024
Tu, W., Deng, W., and Gedeon, T · 2024
Closest in time.
Q-bench: A benchmark for general-purpose foundation models on low-level vision, 2024
Wu, H., Zhang, Z., Zhang, E., Chen, C., Liao, L., Wang, A., Li, C., Sun, W., Yan, Q., Zhai, G., and Lin, W · 2024
Closest in time.
Chartbench: A benchmark for complex visual reasoning in charts, 2024
Xu, Z., Du, S., Qi, Y., Xu, C., Yuan, C., and Guo, J · 2024
Closest in time.
Large language model benchmarks in medical tasks, 2024
Yan, L. K. Q., Niu, Q., Li, M., Zhang, Y., Yin, C. H., Fei, C., Peng, B., Bi, Z., Feng, P., Chen, K., Wang, T., Wang, Y., Chen, S., Liu, M., and Liu, J · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
OpenAI · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Cited alongside, same era.
Pmc-vqa: Visual instruction tuning for medical visual question answering
Zhang, X., Wu, C., Zhao, Z., Lin, W., Zhang, Y., Wang, Y., and Xie, W · 2023
Cited alongside, same era.
Barucci, A., Ciacci, G., Liò, P., Azevedo, T., Cencio, A. D., Merella, M., Bianucci, G., Bosio, G., Casati, S., and Collareta, A · 2024
Cited alongside, same era.
Deng, L., Liu, Y., Li, B., Luo, D., Wu, L., Zhang, C., Lyu, P., Zhang, Z., Zhang, G., Ding, E., Zhu, Y., and Bai, X · 2024
Cited alongside, same era.
Blink: Multimodal large language models can see but not perceive, 2024
Fu, X., Hu, Y., Li, B., Feng, Y., Wang, H., Lin, X., Roth, D., Smith, N. A., Ma, W.-C., and Krishna, R · 2024
Cited alongside, same era.
Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024
GLM, T., Zeng, A., Xu, B., Wang, B., Zhang, C., Yin, D., Rojas, D., Feng, G., Zhao, H., Lai, H., Yu, H., Wang, H., Sun, J., Zhang, J., Cheng, J., Gui, J., Tang, J., Zhang, J., Li, J., Zhao, L., Wu, L., Zhong, L., Liu, M., Huang, M., Zhang, P., Zheng, Q., Lu, R., Duan, S., Zhang, S., Cao, S., Yang, S., Tam, W. L., Zhao, W., Liu, X., Xia, X., Zhang, X., Gu, X., Lv, X., Liu, X., Liu, X., Yang, X., Song, X., Zhang, X., An, Y., Xu, Y., Niu, Y., Yang, Y., Li, Y., Bai, Y., Dong, Y., Qi, Z., Wang, Z., Yang, Z., Du, Z., Hou, Z., and Wang, Z · 2024
Cited alongside, same era.
Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., He, Z., Chen, Q., Zhou, H., Zou, Z., Zhang, H., Hu, S., Zheng, Z., Zhou, J., Cai, J., Han, X., Zeng, G., Li, D., Liu, Z., and Sun, M · 2024
Closest in time.
mplug-owl3: Towards long image-sequence understanding in multi-modal large language models, 2024
Ye, J., Xu, H., Liu, H., Hu, A., Yan, M., Qian, Q., Zhang, J., Huang, F., and Zhou, J · 2024
Closest in time.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L · 2024
Closest in time.
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
Yue, X., Ni, Y., Zheng, T., Zhang, K., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2024
Closest in time.
MATHVERSE: does your multi-modal LLM truly see the diagrams in visual math problems?
Zhang, R., Jiang, D., Zhang, Y., Lin, H., Guo, Z., Qiu, P., Zhou, A., Lu, P., Chang, K., Qiao, Y., Gao, P., and Li, H · 2024
Closest in time.
Uniaa: A unified multi-modal image aesthetic assessment baseline and benchmark, 2024
Zhou, Z., Wang, Q., Lin, B., Su, Y., Chen, R., Tao, X., Zheng, A., Yuan, L., Wan, P., and Zhang, D · 2024
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2024
Closest in time.
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., and Lin, J · 2025
Closest in time.
G-LLaVA: Solving geometric problem with multi-modal large language model
Gao, J., Pi, R., Zhang, J., Ye, J., Zhong, W., Wang, Y., HONG, L., Han, J., Xu, H., Li, Z., and Kong, L · 2025
Closest in time.
Ii-bench: An image implication understanding benchmark for multimodal large language models, 2025
Liu, Z., Fang, F., Feng, X., Du, X., Zhang, C., Wang, Z., Bai, Y., Zhao, Q., Fan, L., Gan, C., Lin, H., Li, J., Ni, Y., Wu, H., Narsupalli, Y., Zheng, Z., Li, C., Hu, X., Xu, R., Chen, X., Yang, M., Liu, J., Liu, R., Huang, W., Zhang, G., and Ni, S · 2025
Closest in time.
Introducing gpt -
OpenAI · 2025
Closest in time.
Team, Q · 2025
Closest in time.
Law of vision representation in MLLMs
Yang, S., Zhai, B., You, Q., Yuan, J., Yang, H., and Xu, C · 2025
Closest in time.
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models
Zhu, J., Wang, W., Chen, Z., Liu, Z., Ye, S., Gu, L., Tian, H., Duan, Y., Su, W., Shao, J., Gao, Z., Cui, E., Wang, X., Cao, Y., Liu, Y., Wei, X., Zhang, H., Wang, H., Xu, W., Li, H., Wang, J., Deng, N., Li, S., He, Y., Jiang, T., Luo, J., Wang, Y., He, C., Shi, B., Zhang, X., Shao, W., He, J., Xiong, Y., Qu, W., Sun, P., Jiao, P., Lv, H., Wu, L., Zhang, K., Deng, H., Ge, J., Chen, K., Wang, L., Dou, M., Lu, L., Zhu, X., Lu, T., Lin, D., Qiao, Y., Dai, J., and Wang, W · 2025
Closest in time.
Math-puma: Progressive upward multimodal alignment to enhance mathematical reasoning
Zhuang, W., Huang, X., Zhang, X., and Zeng, J · 2025
Closest in time.
Medxpertqa: Benchmarking expert-level medical reasoning and understanding
Zuo, Y., Qu, S., Li, Y., Chen, Z., Zhu, X., Hua, E., Zhang, K., Ding, N., and Zhou, B · 2025
Closest in time.
Qwen3.5: Towards native multimodal agents
Team, Q · 2026
Closest in time.
Fossil image identification using deep learning ensembles of data augmented multiviews
Hou, C., Lin, X., Huang, H., Xu, S., Fan, J., Shi, Y., and Lv, H · 2041
Closest in time.