Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) such as CLIP are trained via contrastive learning between text and image pairs, resulting in aligned image and text embeddings that are useful for many downstream tasks.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images.(2009), 2009
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba · 2010
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Learning deep representations of fine-grained visual descriptions
Scott Reed, Zeynep Akata, Honglak Lee, and Bernt Schiele · 2016
Earlier work this paper cites.
Learning to describe differences between pairs of similar images
Harsh Jhamtani and Taylor Berg-Kirkpatrick · 2018
Earlier work this paper cites.
Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly
Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata · 2018
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth · 2019
Earlier work this paper cites.
Robust change captioning
Dong Huk Park, Trevor Darrell, and Anna Rohrbach · 2019
Earlier work this paper cites.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav · 2019
Earlier work this paper cites.
e-snli-ve: Corrected visual-textual entailment with natural language explanations
Virginie Do, Oana-Maria Camburu, Zeynep Akata, and Thomas Lukasiewicz · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Finding it at another side: A viewpoint-adapted matching encoder for change captioning
Xiangxi Shi, Xu Yang, Jiuxiang Gu, Shafiq Joty, and Jianfei Cai · 2020
Earlier work this paper cites.
Generalizing from a few examples: A survey on few-shot learning
Yaqing Wang, Quanming Yao, James T. Kwok, and Lionel M. Ni · 2020
Earlier work this paper cites.
Towards interpretable natural language understanding with explanations as latent variables
Wangchunshu Zhou, Jinyi Hu, Hanlin Zhang, Xiaodan Liang, Maosong Sun, Chenyan Xiong, and Jian Tang · 2020
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Earlier work this paper cites.
Image difference captioning with instance-level fine-grained feature representation
Qingbao Huang, Yu Liang, Jielong Wei, Yi Cai, Hanyu Liang, Ho-fung Leung, and Qing Li · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Cited alongside, same era.
Clipcap: Clip prefix for image captioning
Ron Mokady, Amir Hertz, and Amit H Bermano · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Promptsource: An integrated development environment and repository for natural language prompts
Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, et al · 2022
Cited alongside, same era.
A clip-hitchhiker’s guide to long video retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2022
Zero-shot robustification of zero-shot models with foundation models
Dyah Adila, Changho Shin, Linrong Cai, and Frederic Sala · 2023
Later among the works it cites.
Understanding prompt engineering may not require rethinking generalization
Victor Akinwande, Yiding Jiang, Dylan Sam, and J Zico Kolter · 2023
Later among the works it cites.
Going beyond nouns with vision & language models using synthetic data
Paola Cascante-Bonilla, Khaled Shehada, James Seale Smith, Sivan Doveh, Donghyun Kim, Rameswar Panda, Gul Varol, Aude Oliva, Vicente Ordonez, Rogerio Feris, and Leonid Karlinsky · 2023
Later among the works it cites.
Teaching structured vision & language concepts to vision & language models
Sivan Doveh, Assaf Arbelle, Sivan Harary, Eli Schwartz, Roei Herzig, Raja Giryes, Rogerio Feris, Rameswar Panda, Shimon Ullman, and Leonid Karlinsky · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cyclip: Cyclic contrastive language-image pretraining
Shashank Goel, Hritik Bansal, Sumit Bhatia, Ryan Rossi, Vishwa Vinay, and Aditya Grover · 2022
Cited alongside, same era.
Clip4idc: Clip for image difference captioning
Zixin Guo, Tzu-Jui Wang, and Jorma Laaksonen · 2022
Cited alongside, same era.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang · 2022
Cited alongside, same era.
Does clip bind concepts? probing compositionality in large image models
Martha Lewis, Nihal V. Nayak, Qinan Yu, Jack Merullo, and Ellie Pavlick · 2022
Cited alongside, same era.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Victor Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y Zou · 2022
Cited alongside, same era.
Compositional visual generation with composable diffusion models
Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum · 2022
Cited alongside, same era.
Visual classification via description from large language models
Sachit Menon and Carl Vondrick · 2022
Cited alongside, same era.
Lijie Fan, Dilip Krishnan, Phillip Isola, Dina Katabi, and Yonglong Tian · 2023
Later among the works it cites.
Leveraging multiple descriptive features for robust few-shot image learning
Zhili Feng, Anna Bair, and J Zico Kolter · 2023
Later among the works it cites.
Datacomp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al · 2023
Later among the works it cites.
Finetune like you pretrain: Improved finetuning of zero-shot vision models
Sachin Goyal, Ananya Kumar, Sankalp Garg, Zico Kolter, and Aditi Raghunathan · 2023
Later among the works it cites.
Neurocomparatives: Neuro-symbolic distillation of comparative knowledge
Phillip Howard, Junlin Wang, Vasudev Lal, Gadi Singer, Yejin Choi, and Swabha Swayamdipta · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach · 2023
Later among the works it cites.
Learning with explanation constraints
Rattana Pukdee, Dylan Sam, J Zico Kolter, Maria-Florina Balcan, and Pradeep Ravikumar · 2023
Later among the works it cites.
Losses over labels: Weakly supervised learning via direct loss construction
Dylan Sam and J Zico Kolter · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2023
Later among the works it cites.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al · 2023
Later among the works it cites.
Follow-up differential descriptions: Language models resolve ambiguities for image classification
Reza Esfandiarpoor and Stephen Bach · 2024
Closest in time.
Onediff: A generalist model for image difference captioning
Erdong Hu, Longteng Guo, Tongtian Yue, Zijia Zhao, Shuning Xue, and Jing Liu · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Bayesian neural networks with domain knowledge priors
Dylan Sam, Rattana Pukdee, Daniel P Jeong, Yewon Byun, and J Zico Kolter · 2024
Closest in time.