Fetching the paper…
Reading the bibliography…
While multi-modal large language models (MLLMs) have shown significant progress on many popular visual reasoning benchmarks, whether they possess abstract visual reasoning abilities remains an open question.
Parts of recognition
WA Richards et al · 1984
Earlier work this paper cites.
Heading in the rat: Determination by environmental shape
Joan Margules and CR Gallistel · 1988
Earlier work this paper cites.
Principles of object perception
Elizabeth S Spelke · 1990
Earlier work this paper cites.
Modularity and development: The case of spatial reorientation
Linda Hermer and Elizabeth Spelke · 1996
Earlier work this paper cites.
2.5-month-old infants’ reasoning about when objects should and should not be occluded
Andréa Aguiar and Renée Baillargeon · 1999
Earlier work this paper cites.
Large number discrimination in 6-month-old infants
Fei Xu and Elizabeth S Spelke · 2000
Earlier work this paper cites.
Discrimination of large and small numerosities by human infants
Jennifer S Lipton and Elizabeth S Spelke · 2004
Earlier work this paper cites.
Abstract number and arithmetic in preschool children
Hilary Barth, Kristen La Mont, Jennifer Lipton, and Elizabeth S Spelke · 2005
Earlier work this paper cites.
Is there a geometric module for spatial orientation? squaring theory and evidence
Ken Cheng and Nora S Newcombe · 2005
Earlier work this paper cites.
Core knowledge
Elizabeth S Spelke and Katherine D Kinzler · 2007
Earlier work this paper cites.
Comparing machines and humans on a visual categorization test
François Fleuret, Ting Li, Charles Dubout, Emma K Wampler, Steven Yantis, and Donald Geman · 2011
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Local outlier detection reconsidered: a generalized view on locality with applications to spatial, video, and network outlier detection
Erich Schubert, Arthur Zimek, and Hans-Peter Kriegel · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson · 2019
Earlier work this paper cites.
On the measure of intelligence
François Chollet · 2019
Earlier work this paper cites.
Learning to make analogies by contrasting abstract relational structure
Felix Hill, Adam Santoro, David GT Barrett, Ari S Morcos, and Timothy Lillicrap · 2019
Earlier work this paper cites.
Deepiq: A human-inspired ai system for solving iq test problems
Jacek Mańdziuk and Adam Żychowski · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Earlier work this paper cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Earlier work this paper cites.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav · 2019
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Raven: A dataset for relational and analogical visual reasoning
Chi Zhang, Feng Gao, Baoxiong Jia, Yixin Zhu, and Song-Chun Zhu · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Bongard-logo: A new benchmark for human-level concept learning and reasoning
Weili Nie, Zhiding Yu, Lei Mao, Ankit B Patel, Yuke Zhu, and Anima Anandkumar · 2020
Cited alongside, same era.
Self-supervised relational reasoning for representation learning
Massimiliano Patacchiola and Amos J Storkey · 2020
Cited alongside, same era.
Squinting at vqa models: Introspecting vqa models with sub-questions
Ramprasaath R Selvaraju, Purva Tendulkar, Devi Parikh, Eric Horvitz, Marco Tulio Ribeiro, Besmira Nushi, and Ece Kamar · 2020
Cited alongside, same era.
Learning representations that support extrapolation
Brainteaser: Lateral thinking puzzles for large language models
Yifan Jiang, Filip Ilievski, Kaixin Ma, and Zhivar Sourati · 2023
Later among the works it cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Later among the works it cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2023
Later among the works it cites.
A review of emerging research directions in abstract visual reasoning
Mikołaj Małkiński · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taylor Webb, Zachary Dulberg, Steven Frankland, Alexander Petrov, Randall O’Reilly, and Jonathan Cohen · 2020
Cited alongside, same era.
Scale-localized abstract reasoning
Yaniv Benny, Niv Pekar, and Lior Wolf · 2021
Cited alongside, same era.
Stratified rule-aware network for abstract visual reasoning
Sheng Hu, Yuqing Ma, Xianglong Liu, Yanlu Wei, and Shihao Bai · 2021
Cited alongside, same era.
Abstraction and analogy-making in artificial intelligence
Melanie Mitchell · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
How much intelligence is there in artificial intelligence? a 2020 update
Han LJ van der Maas, Lukas Snoek, and Claire E Stevenson · 2021
Cited alongside, same era.
Perception matters: Detecting perception failures of vqa models using metamorphic testing
Yuanyuan Yuan, Shuai Wang, Mingyue Jiang, and Tsong Yueh Chen · 2021
Cited alongside, same era.
Comparing humans, gpt-4, and gpt-4v on abstraction and reasoning tasks
Melanie Mitchell, Alessandro B Palmarini, and Arsenii Kirillovich Moskvichev · 2023
Later among the works it cites.
The conceptarc benchmark: Evaluating understanding and generalization in the arc domain
Arseny Moskvichev, Victor Vikram Odouard, and Melanie Mitchell · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Review of large vision models and visual prompt engineering
Jiaqi Wang, Zhengliang Liu, Lin Zhao, Zihao Wu, Chong Ma, Sigang Yu, Haixing Dai, Qiushi Yang, Yiheng Liu, Songyao Zhang, et al · 2023
Later among the works it cites.
V*: Guided visual search as a core mechanism in multimodal llms
Penghao Wu and Saining Xie · 2023
Later among the works it cites.
Visual cropping improves zero-shot question answering of multimodal large language models
Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski · 2023
Later among the works it cites.
Mmicl: Empowering vision-language model with multi-modal in-context learning
Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma, Kaikai An, Liang Chen, Zixuan Liu, Sheng Wang, Wenjuan Han, and Baobao Chang · 2023
Later among the works it cites.
The curious case of nonverbal abstract reasoning with multi-modal large language models
Kian Ahrabian, Zhivar Sourati, Kexuan Sun, Jiarui Zhang, Yifan Jiang, Fred Morstatter, and Jay Pujara · 2024
Closest in time.
Claude 3, 2024
Anthropic · 2024
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi · 2024
Closest in time.
Webvoyager: Building an end-to-end web agent with large multimodal models, 2024
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie · 2024
Closest in time.
Text-based reasoning about vector graphics
Zhenhailong Wang, Joy Hsu, Xingyao Wang, Kuan-Hao Huang, Manling Li, Jiajun Wu, and Heng Ji · 2024
Closest in time.
Theoretical analysis of the inductive biases in deep convolutional networks
Zihao Wang and Lei Wu · 2024
Closest in time.
Exploring perceptual limitation of multimodal large language models
Jiarui Zhang, Jinyi Hu, Mahyar Khayatkhoei, Filip Ilievski, and Maosong Sun · 2024
Closest in time.
Gpt-4v(ision) is a generalist web agent, if grounded, 2024
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su · 2024
Closest in time.