Fetching the paper…
Reading the bibliography…
Recently, Multimodal Large Language Models (MLLMs) and Vision Language Models (VLMs) have shown great promise in language-guided perceptual tasks such as recognition, segmentation, and object detection.
Wechsler intelligence scale for children , volume 1
David Wechsler and Habuku Kodama · 1949
Earlier work this paper cites.
Measuring intelligence with the culture fair tests
Raymond Bernard Cattell and Alberta KS Cattell · 1960
Earlier work this paper cites.
Children’s performance on a spatial analogies task
Dedre Gentner · 1977
Earlier work this paper cites.
Influence of working memory on adult age differences in matrix reasoning
Timothy A Salthouse · 1993
Earlier work this paper cites.
The factor
Arthur R Jensen · 1998
Earlier work this paper cites.
Raven progressive matrices
Jean Raven · 2003
Earlier work this paper cites.
Enhanced visual processing contributes to matrix reasoning in autism
Isabelle Soulières, Michelle Dawson, Fabienne Samson, Elise B Barbeau, Cherif P Sahyoun, Gary E Strangman, Thomas A Zeffiro, and Laurent Mottron · 2009
Earlier work this paper cites.
The relationship between n-back performance and matrix reasoning—implications for training and transfer
Susanne M Jaeggi, Barbara Studer-Luethi, Martin Buschkuehl, Yi-Fen Su, John Jonides, and Walter J Perrig · 2010
Earlier work this paper cites.
Comparing machines and humans on a visual categorization test
François Fleuret, Ting Li, Charles Dubout, Emma K Wampler, Steven Yantis, and Donald Geman · 2011
Earlier work this paper cites.
Are there really as many neurons in the human brain as stars in the milky way
Bradley Voytek · 2013
Earlier work this paper cites.
Intelligent testing with the WISC-V
Alan S Kaufman, Susan Engi Raiford, and Diane L Coalson · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
The matrix reasoning item bank (mars-ib): novel, open-access abstract reasoning items for adolescents and adults
Gabriele Chierchia, Delia Fuhrmann, Lisa J Knoll, Blanca Piera Pi-Sunyer, Ashok L Sakhardande, and Sarah-Jayne Blakemore · 2019
Earlier work this paper cites.
On the measure of intelligence
François Chollet · 2019
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Raven: A dataset for relational and analogical visual reasoning
Chi Zhang, Feng Gao, Baoxiong Jia, Yixin Zhu, and Song-Chun Zhu · 2019
Earlier work this paper cites.
Scale-localized abstract reasoning
Yaniv Benny, Niv Pekar, and Lior Wolf · 2021
Earlier work this paper cites.
Stratified rule-aware network for abstract visual reasoning
Sheng Hu, Yuqing Ma, Xianglong Liu, Yanlu Wei, and Shihao Bai · 2021
Earlier work this paper cites.
Learning to compose visual relations
Nan Liu, Shuang Li, Yilun Du, Josh Tenenbaum, and Antonio Torralba · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Cited alongside, same era.
Deep learning methods for abstract visual reasoning: A survey on raven’s progressive matrices
Mikołaj Małkiński and Jacek Mańdziuk · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou · 2022
Cited alongside, same era.
A benchmark for compositional visual reasoning
Aimen Zerroug, Mohit Vaishnav, Julien Colin, Sebastian Musslick, and Thomas Serre · 2022
Mental jenga: A counterfactual simulation model of causal judgments about physical support
Liang Zhou, Kevin A Smith, Joshua B Tenenbaum, and Tobias Gerstenberg · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
The curious case of nonverbal abstract reasoning with multi-modal large language models
Kian Ahrabian, Zhivar Sourati, Kexuan Sun, Jiarui Zhang, Yifan Jiang, Fred Morstatter, and Jay Pujara · 2024
Closest in time.
Introducing the next generation of claude
Anthropic · 2024
Closest in time.
An introduction to vision-language modeling
Florian Bordes, Richard Yuanzhe Pang, Anurag Ajay, Alexander C. Li, Adrien Bardes, Suzanne Petryk, Oscar Mañas, Zhiqiu Lin, Anas Mahmoud, Bargav Jayaraman, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Using cognitive psychology to understand gpt-3
Marcel Binz and Eric Schulz · 2023
Cited alongside, same era.
Have we built machines that think like people?
Luca M Schulze Buschoff, Elif Akata, Matthias Bethge, and Eric Schulz · 2023
Cited alongside, same era.
How do large multimodal models really fare in classical vision few-shot challenges? a deep dive
Qing Guo, Prashan Wanigasekara, Jian Zheng, Jacob Zhiyuan Fang, Xinwei Deng, and Chenyang Tao · 2023
Cited alongside, same era.
Visual programming: Compositional visual reasoning without training
Tanmay Gupta and Aniruddha Kembhavi · 2023
Cited alongside, same era.
Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt
Thilo Hagendorff, Sarah Fabi, and Michal Kosinski · 2023
Cited alongside, same era.
Serwan Jassim, Mario Holubar, Annika Richter, Cornelius Wolff, Xenia Ohmer, and Elia Bruni · 2023
Cited alongside, same era.
Cognitive strategies in matrix-reasoning tasks: State of the art
Paulo Guirro Laurence and Elizeu Coutinho Macedo · 2023
Cited alongside, same era.
Closest in time.
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al · 2024
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi · 2024
Closest in time.
Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, et al · 2024
Closest in time.
Language is not all you need: Aligning perception with language models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, et al · 2024
Closest in time.
Marvel: Multidimensional abstraction and reasoning through visual evaluation and learning
Yifan Jiang, Jiarui Zhang, Kexuan Sun, Zhivar Sourati, Kian Ahrabian, Kaixin Ma, Filip Ilievski, and Jay Pujara · 2024
Closest in time.
Seed-bench: Benchmarking multimodal large language models
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
One self-configurable model to solve many abstract visual reasoning problems
Mikołaj Małkiński and Jacek Mańdziuk · 2024
Closest in time.
Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models
Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuofan Zong, Letian Wang, Yu Liu, and Hongsheng Li · 2024
Closest in time.
Testing theory of mind in large language models and humans
James WA Strachan, Dalila Albergo, Giulia Borghini, Oriana Pansardi, Eugenio Scaliti, Saurabh Gupta, Krati Saxena, Alessandro Rufo, Stefano Panzeri, Guido Manzi, et al · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al · 2024
Closest in time.
Llava-o1: Let vision language models reason step-by-step
Guowei Xu, Peng Jin, Li Hao, Yibing Song, Lichao Sun, and Li Yuan · 2024
Closest in time.
Kiva: Kid-inspired visual analogies for testing large multimodal models
Eunice Yiu, Maan Qraitem, Anisa Noor Majhi, Charlie Wong, Yutong Bai, Shiry Ginosar, Alison Gopnik, and Kate Saenko · 2024
Closest in time.
Evaluation of openai o1: Opportunities and challenges of agi
Tianyang Zhong, Zhengliang Liu, Yi Pan, Yutong Zhang, Yifan Zhou, Shizhe Liang, Zihao Wu, Yanjun Lyu, Peng Shu, Xiaowei Yu, et al · 2024
Closest in time.
R1-v: Reinforcing super generalization ability in vision-language models with less than $3
Liang Chen, Lei Li, Haozhe Zhao, Yifan Song, and Vinci · 2025
Closest in time.