Fetching the paper…
Reading the bibliography…
Contrastive vision-language models (VLMs), like CLIP, have gained popularity for their versatile applicability to various downstream tasks.
The MNIST database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Learning a Similarity Metric Discriminatively, with Application to Face Verification
Sumit Chopra, Raia Hadsell, and Yann LeCun · 2005
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images, 2009
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Fine-Grained Visual Comparisons with Local Learning
Aron Yu and Kristen Grauman · 2014
Earlier work this paper cites.
Microsoft COCO Captions: Data Collection and Evaluation Server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick · 2015
Earlier work this paper cites.
Discovering States and Transformations in Image Collections
Phillip Isola, Joseph J. Lim, and Edward H. Adelson · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Improved Deep Metric Learning with Multi-class N-pair Loss Objective
Kihyuk Sohn · 2016
Earlier work this paper cites.
YFCC100M: The New Data in Multimedia Research
Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Leland McInnes, John Healy, and James Melville · 2018
Earlier work this paper cites.
Representation Learning with Contrastive Predictive Coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Partial Multi-Label Learning
Ming-Kun Xie and Sheng-Jun Huang · 2018
Earlier work this paper cites.
Morpho-MNIST: Quantitative Assessment and Diagnostics for Representation Learning
Daniel C. Castro, Jeremy Tan, Bernhard Kainz, Ender Konukoglu, and Ben Glocker · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere
Tongzhou Wang and Phillip Isola · 2020
Earlier work this paper cites.
Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications
Sandhini Agarwal, Gretchen Krueger, Jack Clark, Alec Radford, Jong Wook Kim, and Miles Brundage · 2021
Earlier work this paper cites.
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut · 2021
Earlier work this paper cites.
Partial success in closing the gap between human and machine vision
Robert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel · 2021
Earlier work this paper cites.
Multimodal Neurons in Artificial Neural Networks
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Cited alongside, same era.
OpenCLIP, 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
Cited alongside, same era.
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
COYO-700M: Image-Text Pair Dataset, 2022
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim · 2022
Cited alongside, same era.
Robust multimodal models have outlier features and encode more concepts
Jonathan Crabbé, Pau Rodríguez, Vaishaal Shankar, Luca Zappella, and Arno Blaas · 2023
Later among the works it cites.
Improving CLIP Training with Language Rewrites
Lijie Fan, Dilip Krishnan, Phillip Isola, Dina Katabi, and Yonglong Tian · 2023
Later among the works it cites.
DataComp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al · 2023
Later among the works it cites.
Identifying Implicit Social Biases in Vision-Language Models
Kimia Hamidieh, Haoran Zhang, Thomas Hartvigsen, and Marzyeh Ghassemi · 2023
Later among the works it cites.
Partial multi-label learning with probabilistic graphical disambiguation
Jun-Yi Hang and Min-Ling Zhang · 2023
Later among the works it cites.
Visual Classification via Description from Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Embedding Arithmetic of Multimodal Queries for Image Retrieval
Guillaume Couairon, Matthijs Douze, Matthieu Cord, and Holger Schwenk · 2022
Cited alongside, same era.
Ranking Info Noise Contrastive Estimation: Boosting Contrastive Learning via Ranked Positives
David T. Hoffmann, Nadine Behrmann, Juergen Gall, Thomas Brox, and Mehdi Noroozi · 2022
Cited alongside, same era.
Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning
Victor Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y Zou · 2022
Cited alongside, same era.
A ConvNet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Cited alongside, same era.
X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval
Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan, Ji Zhang, and Rongrong Ji · 2022
Cited alongside, same era.
Disentangling visual and written concepts in CLIP
Joanna Materzyńska, Antonio Torralba, and David Bau · 2022
Cited alongside, same era.
Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP
Thao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh, and Ludwig Schmidt · 2022
Cited alongside, same era.
Sachit Menon and Carl Vondrick · 2023
Later among the works it cites.
Substance or Style: What Does Your Image Embedding Know?
Cyrus Rashtchian, Charles Herrmann, Chun-Sung Ferng, Ayan Chakrabarti, Dilip Krishnan, Deqing Sun, Da-Cheng Juan, and Andrew Tomkins · 2023
Later among the works it cites.
Towards understanding the modality gap in CLIP
Peiyang Shi, Michael C. Welle, Mårten Björkman, and Danica Kragic · 2023
Later among the works it cites.
What does CLIP know about a red circle? Visual prompt engineering for VLMs
Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi · 2023
Later among the works it cites.
Geodesic Multi-Modal Mixup for Robust Fine-Tuning
Junhyuk So, Changdae Oh, Yongtaek Lim, Hoyoon Byun, Minchul Shin, and Kyungwoo Song · 2023
Later among the works it cites.
Propml: probability partial multi-label learning
Łukasz Struski, Adam Pardyl, Jacek Tabor, and Bartosz Zieliński · 2023
Later among the works it cites.
EVA-CLIP: Improved Training Techniques for CLIP at Scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao · 2023
Later among the works it cites.
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
Matthew Trager, Pramuditha Perera, Luca Zancato, Alessandro Achille, Parminder Bhatia, and Stefano Soatto · 2023
Later among the works it cites.
NLLB-CLIP–train performant multilingual image retrieval model on a budget
Alexander Visheratin · 2023
Later among the works it cites.
When are Lemons Purple? The Concept Association Bias of CLIP
Yutaro Yamada, Yingtian Tang, and Ilker Yildirim · 2023
Later among the works it cites.
Sigmoid Loss for Language Image Pre-Training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
Diagnosing and Rectifying Vision Models using Language
Yuhui Zhang, Jeff Z. HaoChen, Shih-Cheng Huang, Kuan-Chieh Wang, James Zou, and Serena Yeung · 2023
Later among the works it cites.
CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable and Controllable Text-Guided Face Manipulation
Chenliang Zhou, Fangcheng Zhong, and Cengiz Öztireli · 2023
Later among the works it cites.
Its Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
Abrar Fahim, Alex Murphy, and Alona Fyshe · 2024
Closest in time.
Does CLIP’s Generalization Performance Mainly Stem from High Train-Test Similarity?
Prasanna Mayilvahanan, Thaddäus Wiedemer, Evgenia Rusak, Matthias Bethge, and Wieland Brendel · 2024
Closest in time.
A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions
Jack Urbanek, Florian Bordes, Pietro Astolfi, Mary Williamson, Vasu Sharma, and Adriana Romero-Soriano · 2024
Closest in time.
Demystifying CLIP Data
Hu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang, Russell Howes, Vasu Sharma, Shang-Wen Li, Gargi Ghosh, Luke Zettlemoyer, and Christoph Feichtenhofer · 2024
Closest in time.