Fetching the paper…
Reading the bibliography…
A fundamental characteristic common to both human vision and natural language is their compositional nature.
Some controversial questions in phonological theory
Noam Chomsky and Morris Halle · 1965
Earlier work this paper cites.
Logics and languages
MJ Cresswell · 1973
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Jerry A Fodor and Zenon W Pylyshyn · 1988
Earlier work this paper cites.
Natural language and natural selection
Steven Pinker and Paul Bloom · 1990
Earlier work this paper cites.
Compositionality
Theo MV Janssen and Barbara H Partee · 1997
Earlier work this paper cites.
Spontaneous evolution of linguistic structure-an iterated learning model of the emergence of regularity and irregularity
Simon Kirby · 2001
Earlier work this paper cites.
Simulated evolution of language: a review of the field
Amy Perfors · 2002
Earlier work this paper cites.
Language evolution: Consensus and controversies
Morten H Christiansen and Simon Kirby · 2003
Earlier work this paper cites.
Iterated learning: A framework for the emergence of language
Kenny Smith, Simon Kirby, and Henry Brighton · 2003
Earlier work this paper cites.
Understanding linguistic evolution by visualizing the emergence of topographic mappings
Henry Brighton and Simon Kirby · 2006
Earlier work this paper cites.
Cumulative cultural evolution in the laboratory: An experimental approach to the origins of structure in human language
Simon Kirby, Hannah Cornish, and Kenny Smith · 2008
Earlier work this paper cites.
Convention: A philosophical study
David Lewis · 2008
Earlier work this paper cites.
Simulating the evolution of language
Angelo Cangelosi and Domenico Parisi · 2012
Earlier work this paper cites.
Communication leads to the emergence of sub-optimal category structures
Catriona Silvey, Simon Kirbey, and Kenny Smith · 2013
Earlier work this paper cites.
Cultural transmission results in convergence towards colour term universals
Jing Xu, Mike Dowman, and Thomas L Griffiths · 2013
Earlier work this paper cites.
From machine learning to machine reasoning
Léon Bottou · 2014
Earlier work this paper cites.
Iterated learning and the evolution of language
Simon Kirby, Tom Griffiths, and Kenny Smith · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton and et al · 2015
Earlier work this paper cites.
Word meanings evolve to selectively preserve distinctions on salient dimensions
Catriona Silvey, Simon Kirby, and Kenny Smith · 2015
Earlier work this paper cites.
Multi-agent cooperation and the emergence of (natural) language
Angeliki Lazaridou, Alexander Peysakhovich, and Marco Baroni · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Andre Martins and Ramon Astudillo · 2016
Earlier work this paper cites.
Iconicity and the emergence of combinatorial structure in language
Tessa Verhoef, Simon Kirby, and Bart De Boer · 2016
Earlier work this paper cites.
The cultural evolution of structured languages in an open-ended, continuous world
Jon W Carr, Kenny Smith, Hannah Cornish, and Simon Kirby · 2017
Cited alongside, same era.
Sequence memory constraints give rise to language-like structure through iterated learning
Hannah Cornish, Rick Dale, Simon Kirby, and Morten H Christiansen · 2017
Cited alongside, same era.
Natural language does not emerge’naturally’in multi-agent dialog
Satwik Kottur, José MF Moura, Stefan Lee, and Dhruv Batra · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Cited alongside, same era.
Emergence of linguistic communication from referential games with symbolic and pixel input
Angeliki Lazaridou, Karl Moritz Hermann, Karl Tuyls, and Stephen Clark · 2018
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Iterated learning for emergent systematicity in vqa
Ankit Vani, Max Schwarzer, Yuchen Lu, Eeshan Dhekane, and Aaron Courville · 2021
Later among the works it cites.
Learning to generate scene graph from natural language supervision
Yiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu, and Yin Li · 2021
Later among the works it cites.
Why is winoground hard? investigating failures in visuolinguistic compositionality
Anuj Diwan, Layne Berry, Eunsol Choi, David Harwath, and Kyle Mahowald · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
Compositionality in animals and humans
Simon W Townsend, Sabrina Engesser, Sabine Stoll, Klaus Zuberbühler, and Balthasar Bickel · 2018
Cited alongside, same era.
Emergence of compositional language with deep generational transmission
Michael Cogswell, Jiasen Lu, Stefan Lee, Devi Parikh, and Dhruv Batra · 2019
Cited alongside, same era.
Shangmin Guo, Yi Ren, Serhii Havrylov, Stella Frank, Ivan Titov, and Kenny Smith · 2019
Cited alongside, same era.
Ease-of-teaching and language structure from emergent communication
Fushan Li and Michael Bowling · 2019
Cited alongside, same era.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
A large-scale study of representation learning with the visual task adaptation benchmark
Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, Andre Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, et al · 2019
Cited alongside, same era.
Later among the works it cites.
Crepe: Can vision-language foundation models reason compositionally?
Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna · 2022
Later among the works it cites.
Multi-label iterated learning for image classification with label ambiguity
Sai Rajeswar, Pau Rodriguez, Soumye Singhal, David Vazquez, and Aaron Courville · 2022
Later among the works it cites.
On the role of population heterogeneity in emergent communication
Mathieu Rita, Florian Strub, Jean-Bastien Grill, Olivier Pietquin, and Emmanuel Dupoux · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou · 2022
Later among the works it cites.
Vl-checklist: Evaluating pre-trained vision-language models with objects, attributes and relations
Tiancheng Zhao, Tianqi Zhang, Mingwei Zhu, Haozhan Shen, Kyusong Lee, Xiaopeng Lu, and Jianwei Yin · 2022
Later among the works it cites.
The reversal curse: Llms trained on" a is b" fail to learn" b is a"
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans · 2023
Later among the works it cites.
Revisiting multimodal representation in contrastive learning: from patch and token embeddings to finite discrete tokens
Yuxiao Chen, Jianbo Yuan, Yu Tian, Shijie Geng, Xinyu Li, Ding Zhou, Dimitris N Metaxas, and Hongxia Yang · 2023
Later among the works it cites.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jian, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D Hwang, et al · 2023
Later among the works it cites.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Later among the works it cites.
Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna · 2023
Later among the works it cites.
Language is not all you need: Aligning perception with language models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, et al · 2023
Later among the works it cites.
Cola: How to adapt vision-language models to compose objects localized with attributes?
Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan A Plummer, Ranjay Krishna, and Kate Saenko · 2023
Later among the works it cites.
When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou · 2023
Later among the works it cites.
Exif as language: Learning cross-modal associations between images and camera metadata
Chenhao Zheng, Ayush Shrivastava, and Andrew Owens · 2023
Later among the works it cites.
Improving compositional generalization using iterated learning and simplicial embeddings
Yi Ren, Samuel Lavoie, Michael Galkin, Danica J Sutherland, and Aaron C Courville · 2024
Closest in time.
Binding touch to everything: Learning unified multimodal tactile representations
Fengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park, Daniel Wang, Yiming Dou, Ziyao Zeng, Xien Chen, Rit Gangopadhyay, Andrew Owens, and Alex Wong · 2024
Closest in time.