Fetching the paper…
Reading the bibliography…
Despite the success of large vision and language models (VLMs) in many downstream applications, it is unclear how well they encode compositional information.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Smart mining for deep metric learning
Ben Harwood, Vijay Kumar BG, Gustavo Carneiro, Ian Reid, and Tom Drummond · 2017
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Matthew Honnibal and Ines Montani · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
A corpus of natural language for visual reasoning
Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi · 2017
Earlier work this paper cites.
Sampling matters in deep embedding learning
Chao-Yuan Wu, R Manmatha, Alexander J Smola, and Philipp Krahenbuhl · 2017
Earlier work this paper cites.
Artificial unintelligence: How computers misunderstand the world
Meredith Broussard · 2018
Earlier work this paper cites.
Deep metric learning with hierarchical triplet loss
Weifeng Ge · 2018
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Approximating CNNs with bag-of-local-features models works surprisingly well on imagenet
Wieland Brendel and Matthias Bethge · 2019
Earlier work this paper cites.
Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger · 2020
Cited alongside, same era.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Cited alongside, same era.
Hard negative mixing for contrastive learning
Yannis Kalantidis, Mert Bulent Sariyildiz, Noe Pion, Philippe Weinzaepfel, and Diane Larlus · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush · 2020
Cited alongside, same era.
About face: A survey of facial recognition evaluation
Inioluwa Deborah Raji and Genevieve Fried · 2021
Later among the works it cites.
Contrastive learning with hard negative samples
Joshua Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka · 2021
Later among the works it cites.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joëlle Pineau, Adina Williams, and Douwe Kiela · 2021
Later among the works it cites.
A fistful of words: Learning transferable visual models from bag-of-words supervision
Ajinkya Tejankar, Bichen Wu, Saining Xie, Madian Khabsa, Hamed Pirsiavash, and Hamed Firooz · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Contrastive learning of medical visual representations from paired images and text
Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz · 2020
Cited alongside, same era.
Large image datasets: A pyrrhic win for computer vision?
Abeba Birhane and Vinay Uday Prabhu · 2021
Cited alongside, same era.
Covr: A test-bed for visually grounded compositional generalization with real images
Ben Bogin, Shivanshu Gupta, Matt Gardner, and Jonathan Berant · 2021
Cited alongside, same era.
Vision-and-language or vision-for-language? On cross-modal influence in multimodal transformers
Stella Frank, Emanuele Bugliarello, and Desmond Elliott · 2021
Cited alongside, same era.
How effective is bert without word ordering? implications for language understanding and data privacy
Jack Hessel and Alexandra Schofield · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi · 2021
Cited alongside, same era.
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al · 2022
Closest in time.
How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?
Hritik Bansal, Da Yin, Masoud Monajatipoor, and Kai-Wei Chang · 2022
Closest in time.
Easily accessible text-to-image generation amplifies demographic stereotypes at large scale, 2022
Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan · 2022
Closest in time.
DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generative Transformers
Jaemin Cho, Abhaysinh Zala, and Mohit Bansal · 2022
Closest in time.
Testing relational understanding in text-guided image generation
Colin Conwell and Tomer D Ullman · 2022
Closest in time.
Why is winoground hard? investigating failures in visuolinguistic compositionality
Anuj Diwan, Layne Berry, Eunsol Choi, David Harwath, and Kyle Mahowald · 2022
Closest in time.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi · 2022
Closest in time.
DALL·E Now Available Without Waitlist — openai.com
OpenAI · 2022
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al · 2022
Closest in time.
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Closest in time.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Closest in time.
Robert Wolfe, Yiwei Yang, Bill Howe, and Aylin Caliskan · 2022
Closest in time.
A benchmark for compositional visual reasoning
Aimen Zerroug, Mohit Vaishnav, Julien Colin, Sebastian Musslick, and Thomas Serre · 2022
Closest in time.
Lit: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer · 2022
Closest in time.