Fetching the paper…
Reading the bibliography…
Can Multimodal Large Language Models (MLLMs) develop an intuitive number sense similar to humans? Targeting this problem, we introduce Visual Number Benchmark (VisNumBench) to evaluate the number sense abilities of MLLMs across a wide range of visual numerical tasks.
Core systems of number
Lisa Feigenson, Stanislas Dehaene, and Elizabeth Spelke · 2004
Earlier work this paper cites.
Morph: A longitudinal image database of normal adult age-progression
Karl Ricanek and Tamirat Tesafaye · 2006
Earlier work this paper cites.
Dating historical color images
Frank Palermo, James Hays, and Alexei A Efros · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
An image is worth more than a thousand favorites: Surfacing the hidden beauty of flickr pictures
Rossano Schifanella, Miriam Redi, and Luca Maria Aiello · 2015
Earlier work this paper cites.
Single-image crowd counting via multi-column convolutional neural network
Yingying Zhang, Desen Zhou, Siqin Chen, Shenghua Gao, and Yi Ma · 2016
Earlier work this paper cites.
Tartanair: A dataset to push the limits of visual slam
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Sebastian Scherer · 2020
Earlier work this paper cites.
Inter-GPS: Interpretable geometry problem solving with formal language and symbolic reasoning
Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-Chun Zhu · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
CLEVR-Math: A dataset for compositional language, visual and mathematical reasoning
Adam Dahlgren Lindström and Savitha Sam Abraham · 2022
Earlier work this paper cites.
Ordinalclip: Learning rank prompts for language-guided ordinal regression
Wanhua Li, Xiaoke Huang, Zheng Zhu, Yansong Tang, Xiu Li, Jie Zhou, and Jiwen Lu · 2022
Earlier work this paper cites.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Xuan Long Do, Jia Qing Tan, Shafiq Joty, and Enamul Hoque · 2022
Earlier work this paper cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, et al · 2023
Earlier work this paper cites.
UniChart: A universal vision-language pretrained model for chart comprehension and reasoning
Ahmed Masry, Parsa Kavehzadeh, Xuan Long Do, Enamul Hoque, and Shafiq Joty · 2023
Cited alongside, same era.
GPT-4V(ision) system card, 2023
OpenAI · 2023
Cited alongside, same era.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang · 2023
Cited alongside, same era.
LLaVAR: Enhanced visual instruction tuning for text-rich image understanding
Yanzhe Zhang, Ruiyi Zhang, Jiuxiang Gu, Yufan Zhou, Nedim Lipka, Diyi Yang, and Tong Sun · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Our next-generation model: Gemini 1.5. ai, 2024
Sundar Pichai and Demis Hassabis · 2024
Later among the works it cites.
Math-llava: Bootstrapping mathematical reasoning for multimodal large language models
Wenhao Shi, Zhiqiang Hu, Yi Bin, Junhua Liu, Yang Yang, See-Kiong Ng, Lidong Bing, and Roy Ka-Wei Lee · 2024
Later among the works it cites.
Llava-o1: Let vision language models reason step-by-step
Guowei Xu, Peng Jin, Li Hao, Yibing Song, Lichao Sun, and Li Yuan · 2024
Later among the works it cites.
xgen-mm (blip-3): A family of open large multimodal models
Le Xue, Manli Shu, Anas Awadalla, Jun Wang, An Yan, Senthil Purushwalkam, Honglu Zhou, Viraj Prabhu, Yutong Dai, Michael S Ryoo, et al · 2024
Later among the works it cites.
Thinking in space: How multimodal large language models see, remember, and recall spaces
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al · 2024
Cited alongside, same era.
Claude 3.5 sonnet model card addendum
AI Anthropic · 2024
Cited alongside, same era.
Evaluating large vision-and-language models on children’s mathematical olympiads
Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Joanna Matthiesen, Kevin Smith, and Joshua B Tenenbaum · 2024
Cited alongside, same era.
Teach clip to develop a number sense for ordinal regression
Yao Du, Qiang Zhai, Weihang Dai, and Xiaomeng Li · 2024
Cited alongside, same era.
Meng Fang, Xiangpeng Wan, Fei Lu, Fei Xing, and Kai Zou · 2024
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Cited alongside, same era.
Ryo Kamoi, Yusen Zhang, Sarkar Snigdha Sarathi Das, Ranran Haoran Zhang, and Rui Zhang · 2024
Cited alongside, same era.
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao · 2024
Cited alongside, same era.
Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie · 2024
Later among the works it cites.
Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025
Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, and Chong Ruan · 2025
Closest in time.
Blink: Multimodal large language models can see but not perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna · 2025
Closest in time.
Google images
Google · 2025
Closest in time.
Emobench-m: Benchmarking emotional intelligence for multimodal large language models
He Hu, Yucheng Zhou, Lianzhong You, Hongbo Xu, Qianning Wang, Zheng Lian, Fei Richard Yu, Fei Ma, and Laizhong Cui · 2025
Closest in time.
Llama 3.2 vision instruct (11b)
Meta AI · 2025
Closest in time.
Qwen2.5-vl, 2025
Qwen Team · 2025
Closest in time.
Desktop wallpapers hd, free desktop backgrounds
WallpapersCraft · 2025
Closest in time.