Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) demonstrate remarkable fluency in understanding visual scenes, yet they exhibit a critical lack in a fundamental cognitive skill: object counting.
Learning to count objects in images
Victor Lempitsky and Andrew Zisserman · 2010
Earlier work this paper cites.
You only look once: Unified, real-time object detection, 2016
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks, 2016
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2016
Earlier work this paper cites.
Single-image crowd counting via multi-column convolutional neural network
Yingying Zhang, Desen Zhou, Siqin Chen, Shenghua Gao, and Yi Ma · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Composition loss for counting, density map estimation and localization in dense crowds, 2018
Haroon Idrees, Muhmmad Tayyab, Kishan Athrey, Dong Zhang, Somaya Al-Maadeed, Nasir Rajpoot, and Mubarak Shah · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering, 2019
Drew A. Hudson and Christopher D. Manning · 2019
Earlier work this paper cites.
Jhu-crowd++: Large-scale crowd counting dataset and a benchmark method, 2020
Vishwanath A. Sindagi, Rajeev Yasarla, and Vishal M. Patel · 2020
Earlier work this paper cites.
Generative deep-neural-network mixture modeling with semi-supervised minmax+em learning
Nilay Pande and Suyash P. Awate · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Nwpu-crowd: A large-scale benchmark for crowd counting and localization
Qi Wang, Junyu Gao, Wei Lin, and Xuelong Li · 2021
Earlier work this paper cites.
Maea: Multimodal attribution for embodied ai
Vidhi Jain, Jayant Sravan Tamarapalli, Sahiti Yerramilli, and Yonatan Bisk · 2023
Earlier work this paper cites.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Earlier work this paper cites.
Segment anything, 2023
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick · 2023
Earlier work this paper cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Cited alongside, same era.
Training-free object counting with prompts, 2023
Zenglin Shi, Ying Sun, and Mengmi Zhang · 2023
Cited alongside, same era.
Embodied symbiotic assistants that see, act, infer and chat
Carnegie Mellon University · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training, 2023
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Cited alongside, same era.
Pixtral 12b, 2024
Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Baptiste Bout, Devendra Chaplot, Jessica Chudnovsky, Diogo Costa, Baudouin De Monicault, Saurabh Garg, Theophile Gervet, Soham Ghosh, Amélie Héliou, Paul Jacob, Albert Q. Jiang, Kartik Khandelwal, Timothée Lacroix, Guillaume Lample, Diego Las Casas, Thibaut Lavril, Teven Le Scao, Andy Lo, William Marshall, Louis Martin, Arthur Mensch, Pavankumar Muddireddy, Valera Nemychnikova, Marie Pellat, Patrick Von Platen, Nikhil Raghuraman, Baptiste Rozière, Alexandre Sablayrolles, Lucile Saulnier, Romain Sauvestre, Wendy Shang, Roman Soletskyi, Lawrence Stewart, Pierre Stock, Joachim Studnia, Sandeep Subramanian, Sagar Vaze, Thomas Wang, and Sophia Yang · 2024
Cited alongside, same era.
Evaluating and advancing multimodal large language models in perception ability lens, 2025
Feng Chen, Chenhui Gou, Jing Liu, Yang Yang, Zhaoyang Li, Jiyuan Zhang, Zhenbang Sun, Bohan Zhuang, and Qi Wu · 2025
Closest in time.
Huemanity: Probing fine-grained visual perception in mllms
Rynaa Grover, Jayant Sravan Tamarapalli, Sahiti Yerramilli, and Nilay Pande · 2025
Closest in time.
Ai guide dog: Egocentric path prediction on smartphone
Aishwarya Jadhav, Jeffery Cao, Abhishree Shetty, Urvashi Kumar, Aditi Sharma, Ben Sukboontip, Jayant Tamarapalli, Jingyi Zhang, and Aniruddh Koul · 2025
Closest in time.
Mllm-compbench: A comparative reasoning benchmark for multimodal llms, 2025
Jihyung Kil, Zheda Mai, Justin Lee, Zihe Wang, Kerrie Cheng, Lemeng Wang, Ye Liu, Arpita Chowdhury, and Wei-Lun Chao · 2025
Closest in time.
Counts: Benchmarking object detectors and multimodal large language models under distribution shifts, 2025
Jiansheng Li, Xingxuan Zhang, Hao Zou, Yige Guo, Renzhe Xu, Yilong Liu, Chuzhao Zhu, Yue He, and Peng Cui · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Seed-bench: Benchmarking multimodal large language models
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan · 2024
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2024
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang · 2024
Cited alongside, same era.
Vibe-eval: A hard evaluation suite for measuring progress of multimodal language models
Piotr Padlewski, Max Bain, Matthew Henderson, Zhongkai Zhu, Nishant Relan, Hai Pham, Donovan Ong, Kaloyan Aleksiev, Aitor Ormazabal, Samuel Phua, Ethan Yeo, Eugenie Lamprecht, Qi Liu, Yuqi Wang, Eric Chen, Deyu Fu, Lei Li, Che Zheng, Cyprien de Masson d’Autume, Dani Yogatama, Mikel Artetxe, and Yi Tay · 2024
Cited alongside, same era.
Sam 2: Segment anything in images and videos, 2024
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, and Christoph Feichtenhofer · 2024
Cited alongside, same era.
A community-centric perspective for characterizing and detecting anti-asian violence-provoking speech, 2024
Gaurav Verma, Rynaa Grover, Jiawei Zhou, Binny Mathew, Jordan Kraemer, Munmun De Choudhury, and Srijan Kumar · 2024
Cited alongside, same era.
Countgd: Multi-modal open-world counting, 2025
Niki Amini-Naieni, Tengda Han, and Andrew Zisserman · 2025
Cited alongside, same era.
Introducing claude 4, 2025
Anthropic · 2025
Cited alongside, same era.
Closest in time.
Announcing pixtral large, 2024
Mistral AI · 2025
Closest in time.
Introducing mistral medium 3, 2025a
Mistral AI · 2025
Closest in time.
Announcing mistral small 3.1, 2025b
Mistral AI · 2025
Closest in time.
Introducing openai o3 and o4-mini, 2025
OpenAI · 2025
Closest in time.
Lvlm-count: Enhancing the counting ability of large vision-language models, 2025
Muhammad Fetrat Qharabagh, Mohammadreza Ghofrani, and Kimon Fountoulakis · 2025
Closest in time.
T2icount: Enhancing cross-modal understanding for zero-shot counting, 2025
Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, and Michael P. Pound · 2025
Closest in time.
Mathglance: Multimodal large language models do not know where to look in mathematical diagrams, 2025
Yanpeng Sun, Shan Zhang, Wei Tang, Aotian Chen, Piotr Koniusz, Kai Zou, Yuan Xue, and Anton van den Hengel · 2025
Closest in time.
Geochain: Multimodal chain-of-thought for geographic reasoning
Sahiti Yerramilli, Nilay Pande, Rynaa Grover, and Jayant Sravan Tamarapalli · 2025
Closest in time.