Fetching the paper…
Reading the bibliography…
Compositional understanding is crucial for human intelligence, yet it remains unclear whether contemporary vision models exhibit it.
The Language of Thought
Fodor, J. A. and Fodor, J. A · 1975
Earlier work this paper cites.
Deep Residual Learning for Image Recognition, 2015
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Discovering states and transformations in image collections
Isola, P., Lim, J. J., and Adelson, E. H · 2015
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
Abstract representations emerge naturally in neural networks trained to perform multiple tasks
Johnston, S. and Fusi, S · 2017
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization, 2017
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
dSprites: Disentanglement testing sprites dataset, 2017
Matthey, L., Higgins, I., Hassabis, D., and Lerchner, A · 2017
Earlier work this paper cites.
Teaching Compositionality to CNNs, 2017
Stone, A., Wang, H., Stark, M., Liu, Y., Phoenix, D. S., and George, D · 2017
Earlier work this paper cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Valle-Pérez, G., Camargo, C. Q., and Louis, A. A · 2018
Earlier work this paper cites.
Measuring Compositionality in Representation Learning, 2019
Andreas, J · 2019
Earlier work this paper cites.
On the transfer of inductive bias from simulation to the real world: a new disentanglement dataset
Gondal, M. W., Wuthrich, M., Miladinovic, D., Locatello, F., Breidt, M., Volchkov, V., Akpo, J., Bachem, O., Schölkopf, B., and Bauer, S · 2019
Earlier work this paper cites.
Disentangling by Factorising, 2019
Kim, H. and Mnih, A · 2019
Earlier work this paper cites.
Invariant Risk Minimization, 2020
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 2020
Earlier work this paper cites.
A causal view of compositional zero-shot recognition, 2020
Atzmon, Y., Kreuk, F., Shalit, U., and Chechik, G · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
On the transfer of disentangled representations in realistic settings
Dittadi, A., Träuble, F., Locatello, F., Wüthrich, M., Agrawal, V., Winther, O., Bauer, S., and Schölkopf, B · 2020
Earlier work this paper cites.
Shortcut Learning in Deep Neural Networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Earlier work this paper cites.
In Search of Lost Domain Generalization, 2020
Gulrajani, I. and Lopez-Paz, D · 2020
Earlier work this paper cites.
Scaling Laws for Neural Language Models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Object-Centric Learning with Slot Attention, 2020
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Earlier work this paper cites.
The role of Disentanglement in Generalisation
Montero, M. L., Ludwig, C. J., Costa, R. P., Malhotra, G., and Bowers, J · 2020
Earlier work this paper cites.
Compositional Languages Emerge in a Neural Iterated Learning Model, 2020
Ren, Y., Guo, S., Labeau, M., Cohen, S. B., and Kirby, S · 2020
Earlier work this paper cites.
Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization, 2020
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P · 2020
Earlier work this paper cites.
Zero-Shot Learning – A Comprehensive Evaluation of the Good, the Bad and the Ugly, 2020
Xian, Y., Lampert, C. H., Schiele, B., and Akata, Z · 2020
Earlier work this paper cites.
Emerging Properties in Self-Supervised Vision Transformers, 2021
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, 2021
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Earlier work this paper cites.
When and how CNNs generalize to out-of-distribution category-viewpoint combinations, 2021
Madan, S., Henry, T., Dozier, J., Ho, H., Bhandari, N., Sasaki, T., Durand, F., Pfister, H., and Boix, X · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs, 2021
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A · 2021
Cited alongside, same era.
Symbols and mental programs: A hypothesis about human singularity
Dehaene, S., Al Roumi, F., Lakretz, Y., Planton, S., and Sablé-Meyer, M · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Data factors for better compositional generalization
Zhou, C., Chen, P., Liu, B., Li, X., Zhang, C., and Huang, H · 2023
Later among the works it cites.
Don’t cut corners: Exact conditions for modularity in biologically inspired representations
Dorrell, W., Hsu, K., Hollingsworth, L., Lee, J. H., Wu, J., Finn, C., Latham, P. E., Behrens, T. E., and Whittington, J. C · 2024
Later among the works it cites.
Compositional Generative Modeling: A Single Model is Not All You Need, 2024
Du, Y. and Kaelbling, L · 2024
Later among the works it cites.
A complexity-based theory of compositionality
Elmoznino, E., Jiralerspong, T., Bengio, Y., and Lajoie, G · 2024
Later among the works it cites.
Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks, 2024
He, T., Doshi, D., Das, A., and Gromov, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lost in Latent Space: Disentangled Models and the Challenge of Combinatorial Generalisation, 2022
Montero, M. L., Bowers, J. S., Costa, R. P., Ludwig, C. J. H., and Malhotra, G · 2022
Cited alongside, same era.
Visual Representation Learning Does Not Generalize Strongly Within the Same Domain, 2022
Schott, L., von Kügelgen, J., Träuble, F., Gehler, P., Russell, C., Bethge, M., Schölkopf, B., Locatello, F., and Brendel, W · 2022
Cited alongside, same era.
Disentanglement with biological constraints: A theory of functional cell types
Whittington, J. C., Dorrell, W., Ganguli, S., and Behrens, T. E · 2022
Cited alongside, same era.
A benchmark for compositional visual reasoning
Zerroug, A., Vaishnav, M., Colin, J., Musslick, S., and Serre, T · 2022
Cited alongside, same era.
A Theory for Emergence of Complex Skills in Language Models, 2023
Arora, S. and Goyal, A · 2023
Cited alongside, same era.
PUG: Photorealistic and Semantically Controllable Synthetic Data for Representation Learning, 2023
Bordes, F., Shekhar, S., Ibrahim, M., Bouchacourt, D., Vincent, P., and Morcos, A. S · 2023
Cited alongside, same era.
Sparks of Artificial General Intelligence: Early experiments with GPT-4, 2023
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y · 2023
Cited alongside, same era.
When does compositional structure yield compositional generalization? a kernel theory, 2024
Lippl, S. and Stachenfeld, K · 2024
Later among the works it cites.
Compositional risk minimization, 2024
Mahajan, D., Pezeshki, M., Arnal, C., Mitliagkas, I., Ahuja, K., and Vincent, P · 2024
Later among the works it cites.
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models, 2024
Mamaghan, A. M. K., Papa, S., Johansson, K. H., Bauer, S., and Dittadi, A · 2024
Later among the works it cites.
DINOv2: Learning Robust Visual Features without Supervision, 2024
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P · 2024
Later among the works it cites.
The Geometry of Categorical and Hierarchical Concepts in Large Language Models, 2024
Park, K., Choe, Y. J., Jiang, Y., and Veitch, V · 2024
Later among the works it cites.
Vision language models are blind, 2024
Rahmanzadehgervi, P., Bolton, L., Taesiri, M. R., and Nguyen, A. T · 2024
Later among the works it cites.
From causal to concept-based representation learning
Rajendran, G., Buchholz, S., Aragam, B., Schölkopf, B., and Ravikumar, P · 2024
Later among the works it cites.
Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics, 2024
Ren, Y. and Sutherland, D. J · 2024
Later among the works it cites.
Towards Compositionality in Concept Learning, 2024
Stein, A., Naik, A., Wu, Y., Naik, M., and Wong, E · 2024
Later among the works it cites.
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs, 2024
Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., and Xie, S · 2024
Later among the works it cites.
No ”Zero-Shot” Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance, 2024
Udandarao, V., Prabhu, A., Ghosh, A., Sharma, Y., Torr, P. H. S., Bibi, A., Albanie, S., and Bethge, M · 2024
Later among the works it cites.
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision, 2024
Vani, A., Nguyen, B., Lavoie, S., Krishna, R., and Courville, A · 2024
Later among the works it cites.
Anticipating Future Object Compositions without Forgetting, 2024
Zahran, Y., Burghouts, G., and Eisma, Y. B · 2024
Later among the works it cites.
Can Models Learn Skill Composition from Examples?, 2024
Zhao, H., Kaur, S., Yu, D., Goyal, A., and Arora, S · 2024
Later among the works it cites.
ArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling, 2024
Zhu, W. Y., Ye, K., Ke, J., Yu, J., Guibas, L., Milanfar, P., and Yang, F · 2024
Later among the works it cites.
Diffusion classifiers understand compositionality, but conditions apply
Jeong, Y., Uselis, A., Oh, S. J., and Rohrbach, A · 2025
Closest in time.
Clip behaves like a bag-of-words model cross-modally but not uni-modally
Koishigarina, D., Uselis, A., and Oh, S. J · 2025
Closest in time.
On the rankability of visual embeddings, 2025
Sonthalia, A., Uselis, A., and Oh, S. J · 2025
Closest in time.
Intermediate layer classifiers for ood generalization, 2025
Uselis, A. and Oh, S. J · 2025
Closest in time.
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
Wiedemer, T., Sharma, Y., Prabhu, A., Bethge, M., and Brendel, W · 2025
Closest in time.