Fetching the paper…
Reading the bibliography…
A fundamental characteristic common to both human vision and natural language is their compositional nature.
Logics and languages
MJ Cresswell · 1973
Earlier work this paper cites.
Wordnet: a lexical database for english
George A Miller · 1995
Earlier work this paper cites.
Compositionality
Theo MV Janssen and Barbara H Partee · 1997
Earlier work this paper cites.
From machine learning to machine reasoning
Léon Bottou · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Neural module networks, 2015
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2015
Earlier work this paper cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D Manning · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning, 2016
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Kihyuk Sohn · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
Scan: Learning hierarchical compositional visual concepts
Irina Higgins, Nicolas Sonnerat, Loic Matthey, Arka Pal, Christopher P Burgess, Matko Bosnjak, Murray Shanahan, Matthew Botvinick, Demis Hassabis, and Alexander Lerchner · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M. Lake, Tomer D. Ullman, Joshua B. Tenenbaum, and Samuel J. Gershman · 2017
Earlier work this paper cites.
FOIL it! find one mismatch between image and language caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Compositional attention networks for machine reasoning, 2018
Drew A. Hudson and Christopher D. Manning · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni · 2018
Earlier work this paper cites.
ATOMIC: an atlas of machine commonsense for if-then reasoning
Maarten Sap, Ronan Le Bras, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A. Smith, and Yejin Choi · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Closure: Assessing systematic generalization of clevr models
Dzmitry Bahdanau, Harm de Vries, Timothy J O’Donnell, Shikhar Murty, Philippe Beaudoin, Yoshua Bengio, and Aaron Courville · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 2019
Earlier work this paper cites.
Evaluating text-to-image matching using binary image selection (bison), 2019
Hexiang Hu, Ishan Misra, and Laurens van der Maaten · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi · 2019
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, and Wei-Ying Ma · 2019
Earlier work this paper cites.
Words aren’t enough, their order matters: On the robustness of grounding visual referring expressions
Arjun Akula, Spandana Gella, Yaser Al-Onaizan, Song-Chun Zhu, and Siva Reddy · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Cops-ref: A new dataset and task on compositional referring expression comprehension, 2020
Zhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K. Wong, and Qi Wu · 2020
Earlier work this paper cites.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 2020
Earlier work this paper cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2020
Earlier work this paper cites.
Measuring compositional generalization: A comprehensive method on realistic data
Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, Dmitry Tsarkov, Xiao Wang, Marc van Zee, and Olivier Bousquet · 2020
Cited alongside, same era.
Picking bert’s brain: Probing for linguistic dependencies in contextualized embeddings using representational similarity analysis
Michael Lepori and R Thomas McCoy · 2020
Cited alongside, same era.
What does BERT with vision look at?
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2020
Cited alongside, same era.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2020
Cited alongside, same era.
A benchmark for systematic generalization in grounded language understanding
Laura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt, and Brenden M. Lake · 2020
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Later among the works it cites.
DirectProbe: Studying representations without classifiers
Yichu Zhou and Vivek Srikumar · 2021
Later among the works it cites.
Lafite: Towards language-free training for text-to-image generation
Yufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu, Jiuxiang Gu, Jinhui Xu, and Tong Sun · 2021
Later among the works it cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Closest in time.
Testing relational understanding in text-guided image generation
Colin Conwell and Tomer Ullman · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
COVR: A test-bed for visually grounded compositional generalization with real images
Ben Bogin, Shivanshu Gupta, Matt Gardner, and Jonathan Berant · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
Low-complexity probing via finding subnetworks
Steven Cao, Victor Sanh, and Alexander M Rush · 2021
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Zero-shot out-of-distribution detection based on the pre-trained model clip, 2021
Sepideh Esmaeilpour, Bing Liu, Eric Robertson, and Lei Shu · 2021
Cited alongside, same era.
Vision-and-language or vision-for-language? on cross-modal influence in multimodal transformers
Stella Frank, Emanuele Bugliarello, and Desmond Elliott · 2021
Cited alongside, same era.
Can foundation models perform zero-shot task specification for robot manipulation?
Yuchen Cui, Scott Niekum, Abhinav Gupta, Vikash Kumar, and Aravind Rajeswaran · 2022
Closest in time.
Multi-modal alignment using representation codebook, 2022
Jiali Duan, Liqun Chen, Son Tran, Jinyu Yang, Yi Xu, Belinda Zeng, and Trishul Chilimbi · 2022
Closest in time.
Measuring compositional consistency for video question answering
Mona Gandhi, Mustafa O. Gul, Eva Prakash, Madeleine Grunde-McLaughlin, Ranjay Krishna, and Maneesh Agrawala · 2022
Closest in time.
Cyclip: Cyclic contrastive language-image pretraining, 2022
Shashank Goel, Hritik Bansal, Sumit Bhatia, Ryan A. Rossi, Vishwa Vinay, and Aditya Grover · 2022
Closest in time.
Language-driven semantic segmentation
Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl · 2022
Closest in time.
Align and prompt: Video-and-language pre-training with entity prompts
Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven C.H. Hoi · 2022
Closest in time.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Closest in time.
UNIMO-2: End-to-end unified vision-language grounded learning
Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang · 2022
Closest in time.
Comprehending and ordering semantics for image captioning
Yehao Li, Yingwei Pan, Ting Yao, and Tao Mei · 2022
Closest in time.
Stylet2i: Toward compositional and high-fidelity text-to-image synthesis
Zhiheng Li, Martin Renqiang Min, Kai Li, and Chenliang Xu · 2022
Closest in time.
Cots: Collaborative two-stream vision-language pre-training model for cross-modal retrieval, 2022
Haoyu Lu, Nanyi Fei, Yuqi Huo, Yizhao Gao, Zhiwu Lu, and Ji-Rong Wen · 2022
Closest in time.
Disentangling visual and written concepts in clip, 2022
Joanna Materzynska, Antonio Torralba, and David Bau · 2022
Closest in time.
Finding structural knowledge in multimodal-BERT
Victor Milewski, Miryam de Lhoneux, and Marie-Francine Moens · 2022
Closest in time.
VALSE: A task-independent benchmark for vision and language models centered on linguistic phenomena
Letitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank, Iacer Calixto, and Albert Gatt · 2022
Closest in time.
Exposing the limits of video-text models through contrast sets
Jae Sung Park, Sheng Shen, Ali Farhadi, Trevor Darrell, Yejin Choi, and Anna Rohrbach · 2022
Closest in time.
On guiding visual attention with language specification
Suzanne Petryk, Lisa Dunlap, Keyan Nasseri, Joseph Gonzalez, Trevor Darrell, and Anna Rohrbach · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
Denseclip: Language-guided dense prediction with context-aware prompting
Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu · 2022
Closest in time.
Probing the role of positional information in vision-language models
Philipp J. Rösch and Jindřich Libovický · 2022
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al · 2022
Closest in time.
How much can CLIP benefit vision-and-language tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer · 2022
Closest in time.
Proposalclip: Unsupervised open-category object proposal generation via exploiting clip cues, 2022
Hengcan Shi, Munawar Hayat, Yicheng Wu, and Jianfei Cai · 2022
Closest in time.
Reclip: A strong zero-shot baseline for referring expression comprehension
Sanjay Subramanian, Will Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, and Anna Rohrbach · 2022
Closest in time.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Closest in time.
Winoground: Probing vision and language models for visio-linguistic compositionality, 2022
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Closest in time.
Do prompt-based models really understand the meaning of their prompts?
Albert Webson and Ellie Pavlick · 2022
Closest in time.
Vision-language pre-training with triple contrastive learning
Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, and Junzhou Huang · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Closest in time.