Fetching the paper…
Reading the bibliography…
Vision and Language Models (VLMs), such as CLIP, have enabled visual recognition of a potentially unlimited set of categories described by text prompts.
“Automated Flower Classification Over a Large Number of Classes,”
Maria-Elena Nilsback and Andrew Zisserman, · 2008
Earlier work this paper cites.
“Imagenet: A Large-scale Hierarchical Image Database,”
Jia Deng et al., · 2009
Earlier work this paper cites.
“SUN Database: Large-scale Scene Recognition from Abbey to Zoo,”
Jianxiong Xiao et al., · 2010
Earlier work this paper cites.
“UCF101: A Dataset of 101 Human Actions Classes from Videos in the Wild,”
Khurram Soomro et al., · 2012
Earlier work this paper cites.
“Fine-grained Visual Classification of Aircraft,”
Subhransu Maji et al., · 2013
Earlier work this paper cites.
“Describing Textures in the Wild,”
Mircea Cimpoi et al., · 2014
Earlier work this paper cites.
“Food-101–Mining Discriminative Components with Random Forests,”
Lukas Bossard et al., · 2014
Earlier work this paper cites.
“Rethinking the Inception Architecture for Computer Vision,”
Christian Szegedy et al., · 2016
Earlier work this paper cites.
“Introducing EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification,”
Patrick Helber et al., · 2018
Earlier work this paper cites.
“Fixmatch: Simplifying Semi-supervised Learning with Consistency and Confidence,”
Kihyuk Sohn et al., · 2020
Earlier work this paper cites.
“Learning Transferable Visual Models from Natural Language Supervision,”
Alec Radford et al., · 2021
Earlier work this paper cites.
“The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization,”
Dan Hendrycks et al., · 2021
Cited alongside, same era.
“Align before Fuse: Vision and Language Representation Learning with Momentum Distillation,”
Junnan Li et al., · 2021
Cited alongside, same era.
“Cloob: Modern Hopfield Networks with InfoLOOB Outperform Clip,”
Andreas Fürst et al., · 2021
Cited alongside, same era.
“Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm,”
Yangguang Li et al., · 2021
Cited alongside, same era.
“Clip-adapter: Better Vision-language Models with Feature Adapters,”
Peng Gao et al., · 2021
“Vision-Language Pre-Training with Triple Contrastive Learning,”
Jinyu Yang et al., · 2022
Later among the works it cites.
Junnan Li et al., · 2022
Later among the works it cites.
“CyCLIP: Cyclic Contrastive Language-Image Pretraining,”
Shashank Goel et al., · 2022
Later among the works it cites.
“PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining,”
Yuting Gao et al., · 2022
Later among the works it cites.
“Unsupervised Prompt Learning for Vision-Language Models,”
Tony Huang et al., · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Supervision Exists Everywhere: A data Efficient Contrastive Language-Image Pre-training Paradigm,”
Yangguang Li et al., · 2021
Cited alongside, same era.
“Slip: Self-supervision Meets Language-image Pre-training,”
Norman Mu et al., · 2021
Cited alongside, same era.
“Filip: Fine-grained Interactive Language-image Pre-training,”
Lewei Yao et al., · 2021
Cited alongside, same era.
“Data Efficient Language-supervised Zero-shot Recognition with Optimal Transport Distillation,”
Bichen Wu et al., · 2021
Cited alongside, same era.
“Learning to Prompt for Vision-Language Models,”
Kaiyang Zhou et al., · 2022
Cited alongside, same era.
“Conditional Prompt Learning for Vision-Language Models,”
Kaiyang Zhou et al., · 2022
Cited alongside, same era.
Later among the works it cites.
“Improving Zero-Shot Models with Label Distribution Priors,”
Jonathan Kahana et al., · 2022
Later among the works it cites.
Wei Lin et al., · 2023
Closest in time.
“LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image Collections,”
M Jehanzeb Mirza et al., · 2023
Closest in time.
“Maple: Multi-modal Prompt Learning,”
Muhammad Uzair Khattak et al., · 2023
Closest in time.
“What does a Platypus Look Like? Generating Customized Prompts for Zero-shot Image Classification,”
Sarah Pratt et al., · 2023
Closest in time.
“DrML: Diagnosing and Rectifying Vision Models using Language,”
Yuhui Zhang et al., · 2023
Closest in time.