Fetching the paper…
Reading the bibliography…
Recent work has documented striking heterogeneity in the performance of state-of-the-art vision language models (VLMs), including both multimodal language models and text-to-image models.
Progressive matrices: A perceptual test of intelligence, individual form
J. C. Raven · 1938
Earlier work this paper cites.
The discrimination of visual number
E. L. Kaufman, M. W. Lord, T. W. Reese, and J. Volkmann · 1949
Earlier work this paper cites.
The magical number seven, plus or minus two: Some limits on our capacity for processing information
G. A. Miller · 1956
Earlier work this paper cites.
A feature-integration theory of attention
A. M. Treisman and G. Gelade · 1980
Earlier work this paper cites.
Subitizing: an analysis of its component processes
G. Mandler and B. J. Shebo · 1982
Earlier work this paper cites.
Illusory conjunctions in the perception of objects
A. Treisman and H. Schmidt · 1982
Earlier work this paper cites.
The topography of ability and learning correlations
R. E. Snow, P. C. Kyllonen, B. Marshalek, et al · 1984
Earlier work this paper cites.
Why are small and large numbers enumerated differently? a limited-capacity preattentive stage in vision
L. M. Trick and Z. W. Pylyshyn · 1994
Earlier work this paper cites.
The correlation theory of brain function
C. Von Der Malsburg · 1994
Earlier work this paper cites.
The temporal dynamics of visual search: evidence for parallel processing in feature and conjunction searches
B. McElree and M. Carrasco · 1999
Earlier work this paper cites.
The binding problem
A. L. Roskies · 1999
Earlier work this paper cites.
Does subitizing reflect numerical estimation?
S. K. Revkin, M. Piazza, V. Izard, L. Cohen, and S. Dehaene · 2008
Earlier work this paper cites.
Analogy and relational reasoning
K. J. Holyoak · 2012
Earlier work this paper cites.
The sparseness of mixed selectivity neurons controls the generalization–discrimination trade-off
O. Barak, M. Rigotti, and S. Fusi · 2013
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Cited alongside, same era.
Measuring abstract reasoning in neural networks
D. Barrett, F. Hill, A. Santoro, A. Morcos, and T. Lillicrap · 2018
Cited alongside, same era.
Blender - a 3D modelling and rendering package
B. O. Community · 2018
Cited alongside, same era.
Compositional attention networks for machine reasoning
D. A. Hudson and C. D. Manning · 2018
Cited alongside, same era.
Monet: Unsupervised scene decomposition and representation
C. P. Burgess, L. Matthey, N. Watters, R. Kabra, I. Higgins, M. Botvinick, and A. Lerchner · 2019
Cited alongside, same era.
Winoground: Probing vision and language models for visio-linguistic compositionality, 2022
T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross · 2022
Later among the works it cites.
Gamr: A guided attention model for (visual) reasoning
M. Vaishnav and T. Serre · 2022
Later among the works it cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Later among the works it cites.
Comparing humans, gpt-4, and gpt-4v on abstraction and reasoning tasks
M. Mitchell, A. B. Palmarini, and A. Moskvichev · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raven: A dataset for relational and analogical visual reasoning
C. Zhang, F. Gao, B. Jia, Y. Zhu, and S.-C. Zhu · 2019
Cited alongside, same era.
On the binding problem in artificial neural networks
K. Greff, S. Van Steenkiste, and J. Schmidhuber · 2020
Cited alongside, same era.
Object-centric learning with slot attention
F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf · 2020
Cited alongside, same era.
Attention over learned object embeddings enables complex visual reasoning
D. Ding, F. Hill, A. Santoro, M. Reynolds, and M. Botvinick · 2021
Cited alongside, same era.
No coincidence, george: Capacity-limits as the curse of compositionality
S. M. Frankland, T. Webb, and J. D. Cohen · 2021
Cited alongside, same era.
Testing relational understanding in text-guided image generation
C. Conwell and T. Ullman · 2022
Cited alongside, same era.
Does clip bind concepts? probing compositionality in large image models
M. Lewis, N. V. Nayak, P. Yu, Q. Yu, J. Merullo, S. H. Bach, and E. Pavlick · 2022
Cited alongside, same era.
S. S. Mondal, T. Webb, and J. D. Cohen · 2023
Later among the works it cites.
On the rational boundedness of cognitive control: Shared versus separated representations
S. Musslick, A. Saxe, A. N. Hoskin, Y. Sagiv, D. Reichman, G. Petri, and J. D. Cohen · 2023
Later among the works it cites.
Solving the binding problem: Assemblies form when neurons enhance their firing rate—they don’t need to oscillate or synchronize
P. R. Roelfsema · 2023
Later among the works it cites.
Emergent analogical reasoning in large language models
T. Webb, K. J. Holyoak, and H. Lu · 2023
Later among the works it cites.
Subobject-level image tokenization, 2024
D. Chen, S. Cahyawijaya, J. Liu, B. Wang, and P. Fung · 2024
Closest in time.
Vision language models are blind
P. Rahmanzadehgervi, L. Bolton, M. R. Taesiri, and A. T. Nguyen · 2024
Closest in time.
Can generative multimodal models count to ten?
S. Rane, A. Ku, J. M. Baldridge, I. Tenney, T. L. Griffiths, and B. Kim · 2024
Closest in time.
Kiva: Kid-inspired visual analogies for testing large multimodal models
E. Yiu, M. Qraitem, C. Wong, A. N. Majhi, Y. Bai, S. Ginosar, A. Gopnik, and K. Saenko · 2024
Closest in time.
Good at captioning, bad at counting: Benchmarking gpt-4v on earth observation data
C. Zhang and S. Wang · 2024
Closest in time.