Fetching the paper…
Reading the bibliography…
This paper investigates visual analogical reasoning in large multimodal models (LMMs) compared to human adults and children.
Raven’s progressive matrices
John C Raven and JH Court · 1938
Earlier work this paper cites.
Pattern Recognition
M. M. Bongard · 1970
Earlier work this paper cites.
Mental rotation of three-dimensional objects
Roger N Shepard and Jacqueline Metzler · 1971
Earlier work this paper cites.
Discrimination of color and pattern novelty in one-month human infants
Allen E Milewski and Einar R Siqueland · 1975
Earlier work this paper cites.
Component processes in analogical reasoning
Robert J Sternberg · 1977
Earlier work this paper cites.
The development of analogical reasoning processes
Robert J Sternberg and Bathsheva Rifkin · 1979
Earlier work this paper cites.
Infant perception of the invariant size of approaching and receding objects
RH Day and BE McKenzie · 1981
Earlier work this paper cites.
Structure-mapping: A theoretical framework for analogy
Dedre Gentner · 1983
Earlier work this paper cites.
What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test
Patricia A Carpenter, Marcel A Just, and Peter Shell · 1990
Earlier work this paper cites.
Size constancy at birth: Newborn infants’ responses to retinal and real size
Alan Slater, Anne Mattock, and Elizabeth Brown · 1990
Earlier work this paper cites.
A direct demonstration of functional specialization in human visual cortex
Semir Zeki, JD Watson, CJ Lueck, Karl J Friston, C Kennard, and RS Frackowiak · 1991
Earlier work this paper cites.
Development of calculation abilities in young children
Susan Cohen Levine, Nancy C Jordan, and Janellen Huttenlocher · 1992
Earlier work this paper cites.
Priming, analogy, and awareness in complex reasoning
Christian D Schunn and Kevin Dunbar · 1996
Earlier work this paper cites.
The purdue visualization of rotations test
George M Bodner and Roland B Guay · 1997
Earlier work this paper cites.
Viewpoint dependence in scene recognition
Vaibhav A Diwadkar and Timothy P McNamara · 1997
Earlier work this paper cites.
Color-based object recognition
Theo Gevers and Arnold WM Smeulders · 1999
Earlier work this paper cites.
The mental cutting test" schnitte" and the picture rotation test-two new measures to assess spatial ability
Claudia Quaiser-Pohl · 2003
Earlier work this paper cites.
The development of visual short-term memory capacity in infants
Shannon Ross-sheehy, Lisa M Oakes, and Steven J Luck · 2003
Earlier work this paper cites.
The triumph of numbers: How counting shaped modern life
I Bernard Cohen · 2005
Earlier work this paper cites.
Why is real-world visual object recognition hard?
Nicolas Pinto, David D Cox, and James J DiCarlo · 2008
Earlier work this paper cites.
Where hypotheses come from: Learning new relations by structural alignment
Stella Christie and Dedre Gentner · 2010
Earlier work this paper cites.
Visual experience: sensation, cognition, and constancy
Gary Hatfield and Sarah Allred · 2012
Earlier work this paper cites.
Analogy and relational reasoning
Keith J Holyoak · 2012
Earlier work this paper cites.
Cross-cultural differences in cognitive development: Attention to relations and objects
Megumi Kuwabara and Linda B Smith · 2012
Earlier work this paper cites.
Development of mental rotation in 3-to 5-year-old children
Andrea Frick, Melissa A Hansen, and Nora S Newcombe · 2013
Earlier work this paper cites.
Analogical reasoning in children
Usha Goswami · 2013
Earlier work this paper cites.
Understanding spatial transformations: Similarities and differences between mental rotation and mental folding
Justin Harris, Kathy Hirsh-Pasek, and Nora S Newcombe · 2013
Earlier work this paper cites.
Infants detect changes in everyday scenes: The role of scene gist
Shinchieh Duh and Su-hua Wang · 2014
Earlier work this paper cites.
Correlation of motor skill, mental rotation, and working memory in 3-to 6-year-old children
Jennifer Lehmann, Claudia Quaiser-Pohl, and Petra Jansen · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
jspsych: A javascript library for creating behavioral experiments in a web browser
Joshua R De Leeuw · 2015
Earlier work this paper cites.
Single-image depth perception in the wild
Weifeng Chen, Zhao Fu, Dawei Yang, and Jia Deng · 2016
Cited alongside, same era.
Infants actively construct and update their representations of physical events: Evidence from change detection by 12-month-olds
Su-hua Wang and Elizabeth J Goldman · 2016
Cited alongside, same era.
Counting everyday objects in everyday scenes
Prithvijit Chattopadhyay, Ramakrishna Vedantam, Ramprasaath R Selvaraju, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Modeling visual problem solving as analogical reasoning
Andrew Lovett and Kenneth Forbus · 2017
Cited alongside, same era.
Navigating without vision: Principles of blind spatial cognition
The development of color perception and cognition
John Maule, Alice E Skelton, and Anna Franklin · 2023
Later among the works it cites.
Comparing humans, gpt-4, and gpt-4v on abstraction and reasoning tasks
Melanie Mitchell, Alessandro B Palmarini, and Arseny Moskvichev · 2023
Later among the works it cites.
The conceptarc benchmark: Evaluating understanding and generalization in the arc domain
Arseny Moskvichev, Victor Vikram Odouard, and Melanie Mitchell · 2023
Later among the works it cites.
Gpt-4v(ision) technical work and authors
OpenAI · 2023
Later among the works it cites.
Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicholas A Giudice · 2018
Cited alongside, same era.
On the measure of intelligence
François Chollet · 2019
Cited alongside, same era.
What object should i use?-task driven object detection
Johann Sawatzky, Yaser Souri, Christian Grund, and Jurgen Gall · 2019
Cited alongside, same era.
Transformations and transfer: Preschool children understand abstract relations and reason analogically in a causal task
Mariel K Goddu, Tania Lombrozo, and Alison Gopnik · 2020
Cited alongside, same era.
Visual chirality
Zhiqiu Lin, Jin Sun, Abe Davis, and Noah Snavely · 2020
Cited alongside, same era.
Visual size processing in early visual cortex follows lateral occipital cortex involvement
Hang Zeng, Gereon R Fink, and Ralph Weidner · 2020
Cited alongside, same era.
Can language models encode perceptual structure without grounding? a case study in color
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard · 2021
Cited alongside, same era.
Perception test: A diagnostic benchmark for multimodal video models
Viorica Pătrăucean, Lucas Smaira, Ankush Gupta, Adrià Recasens Continente, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Joseph Heyward, Mateusz Malinowski, Yi Yang, et al · 2023
Later among the works it cites.
Molly R Petersen and Lonneke van der Plas · 2023
Later among the works it cites.
Can large language models really improve by self-critiquing their own plans?
Karthik Valmeekam, Matthew Marquez, and Subbarao Kambhampati · 2023
Later among the works it cites.
Images speak in images: A generalist painter for in-context visual learning
Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, and Tiejun Huang · 2023
Later among the works it cites.
Emergent analogical reasoning in large language models
Taylor Webb, Keith J Holyoak, and Hongjing Lu · 2023
Later among the works it cites.
Perception and simulation during concept learning
Erik Weitnauer, Robert L Goldstone, and Helge Ritter · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang · 2023
Later among the works it cites.
The curious case of nonverbal abstract reasoning with multi-modal large language models
Kian Ahrabian, Zhivar Sourati, Kexuan Sun, Jiarui Zhang, Yifan Jiang, Fred Morstatter, and Jay Pujara · 2024
Closest in time.
Sequential modeling enables scalable learning for large vision models
Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan L Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros · 2024
Closest in time.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter, Danny Driess, Pete Florence, Dorsa Sadigh, Leonidas Guibas, and Fei Xia · 2024
Closest in time.
Solving bongard problems with a visual language and pragmatic constraints
Stefan Depeweg, Contantin A Rothkopf, and Frank Jäkel · 2024
Closest in time.
Blink: Multimodal large language models can see but not perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna · 2024
Closest in time.
Language is not all you need: Aligning perception with language models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, et al · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Openai o1 system card
OpenAI · 2024
Closest in time.
Prolific: Online platform for participant recruitment
Prolific · 2024
Closest in time.
Vision language models are blind
Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri, and Anh Totti Nguyen · 2024
Closest in time.
Understanding llms requires more than statistical generalization
Patrik Reizinger, Szilvia Ujváry, Anna Mészáros, Anna Kerekes, Wieland Brendel, and Ferenc Huszár · 2024
Closest in time.
Colorfoil: Investigating color blindness in large vision and language models
Ahnaf Mozib Samin, M Firoz Ahmed, and Md Mushtaq Shahriyar Rafee · 2024
Closest in time.
Children helping science: The online directory of research studies for children
Children Helping Science · 2024
Closest in time.
A vision check-up for language models
Pratyusha Sharma, Tamar Rott Shaham, Manel Baradad, Stephanie Fu, Adrian Rodriguez-Munoz, Shivam Duggal, Phillip Isola, and Antonio Torralba · 2024
Closest in time.
Is a picture worth a thousand words? delving into spatial reasoning for vision language models
Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, Yixuan Li, and Neel Joshi · 2024
Closest in time.
A surprising failure? multimodal llms and the nlvr challenge
Anne Wu, Kianté Brantley, and Yoav Artzi · 2024
Closest in time.
Understanding the limits of vision language models through the lens of the binding problem
Declan Campbell, Sunayana Rane, Tyler Giallanza, Camillo Nicolò De Sabbata, Kia Ghods, Amogh Joshi, Alexander Ku, Steven Frankland, Tom Griffiths, Jonathan D Cohen, et al · 2025
Closest in time.
Causal relational problem solving in toddlers
Mariel K Goddu, Eunice Yiu, and Alison Gopnik · 2025
Closest in time.