Fetching the paper…
Reading the bibliography…
Current multimodal benchmarks often conflate reasoning with domain-specific knowledge, making it difficult to isolate and evaluate general reasoning abilities in non-expert settings.
Differential involvement of left prefrontal cortexin inductive and deductive reasoning
Vinod Goel and Raymond J Dolan · 2004
Earlier work this paper cites.
Connecting long distance: semantic distance in analogical reasoning modulates frontopolar cortex activity
Adam E Green, David JM Kraemer, Jonathan A Fugelsang, Jeremy R Gray, and Kevin N Dunbar · 2010
Earlier work this paper cites.
Causal knowledge and the development of inductive reasoning
Aimée K Bright and Aidan Feeney · 2014
Earlier work this paper cites.
The interaction of process and domain in prefrontal cortex during inductive reasoning
Laura Babcock and Antonino Vallesi · 2015
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning, 2020
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang · 2020
Earlier work this paper cites.
Claude, 2022
https://www.anthropic.com/index/introducing-claude Anthropic · 2022
Earlier work this paper cites.
Vasr: Visual analogies of situation recognition
Yonatan Bitton, Ron Yosef, Eliyahu Strugo, Dafna Shahaf, Roy Schwartz, and Gabriel Stanovsky · 2023
Earlier work this paper cites.
Are deep neural networks smarter than second graders?
Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Kevin A. Smith, and Joshua B. Tenenbaum · 2023
Earlier work this paper cites.
Lora: A logical reasoning augmented dataset for visual question answering
Jingying Gao, Qi Wu, Alan Blair, and Maurice Pagnucco · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2023
Cited alongside, same era.
Qwen-VL: A versatile vision-language model for understanding, localization, text reading, and beyond, 2024
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2024
Cited alongside, same era.
Hello gpt4-o. https://openai.com/index/hello-gpt-4o/, 2024
OpenAI · 2024
Later among the works it cites.
Qvq: To see the world with wisdom, December 2024
Qwen Team · 2024
Later among the works it cites.
SciFIBench: Benchmarking large multimodal models for scientific figure interpretation
Jonathan Roberts, Kai Han, Neil Houlsby, and Samuel Albanie · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al · 2024
Later among the works it cites.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Cited alongside, same era.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Cited alongside, same era.
Lmms-eval: Accelerating the development of large multimoal models, March 2024
Bo Li*, Peiyuan Zhang*, Kaicheng Zhang*, Fanyi Pu*, Xinrun Du, Yuhao Dong, Haotian Liu, Yuanhan Zhang, Ge Zhang, Chunyuan Li, and Ziwei Liu · 2024
Cited alongside, same era.
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li · 2024
Cited alongside, same era.
Are deep neural networks smarter than second graders?
Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Kevin Smith, and Joshua B Tenenbaum
Cited in the paper.
Improved baselines with visual instruction tuning, 2023a
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee
Cited in the paper.
Llava-next: Improved reasoning, ocr, and world knowledge, 2024a
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee
Cited in the paper.
Harnessing webpage uis for text-rich visual understanding
Junpeng Liu, Tianyue Ou, Yifan Song, Yuxiao Qu, Wai Lam, Chenyan Xiong, Wenhu Chen, Graham Neubig, and Xiang Yue
Cited in the paper.
Humanity’s Last Exam’s Authors · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
Pangea: A fully open multilingual multimodal LLM for 39 languages
Xiang Yue, Yueqi Song, Akari Asai, Simran Khanuja, Anjali Kantharuban, Seungone Kim, Jean de Dieu Nyandwi, Lintang Sutawika, Sathyanarayanan Ramamoorthy, and Graham Neubig · 2025
Closest in time.