Fetching the paper…
Reading the bibliography…
Foundational Vision-Language models such as CLIP have exhibited impressive generalization in downstream tasks.
A mathematical theory of evidence
Glenn Shafer · 1976
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Li Fei-Fei, Robert Fergus, and Pietro Perona · 2007
Earlier work this paper cites.
Upper and lower probabilities induced by a multivalued mapping
Arthur P Dempster · 2008
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Causal inference in statistics: An overview
Judea Pearl · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
SUN database: Large-scale scene recognition from abbey to zoo
Jianxiong Xiao, James Hays, Krista A. Ehinger, Aude Oliva, and Antonio Torralba · 2010
Earlier work this paper cites.
Cats and dogs
Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew B. Blaschko, and Andrea Vedaldi · 2013
Earlier work this paper cites.
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
Food-101 - mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Earlier work this paper cites.
Causal inference in statistics: A primer
Madelyn Glymour, Judea Pearl, and Nicholas P Jewell · 2016
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Earlier work this paper cites.
Learning robust global representations by penalizing local predictive power
Haohan Wang, Songwei Ge, Zachary C. Lipton, and Eric P. Xing · 2019
Cited alongside, same era.
Counterfactual samples synthesizing for robust visual question answering
Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang · 2020
Cited alongside, same era.
Interventional few-shot learning
Zhongqi Yue, Hanwang Zhang, Qianru Sun, and Xian-Sheng Hua · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
The role of deconfounding in meta-learning
Yinjie Jiang, Zhengyu Chen, Kun Kuang, Luotian Yuan, Xinhai Ye, Zhihua Wang, Fei Wu, and Ying Wei · 2022
Later among the works it cites.
Interventional contrastive learning with meta semantic regularizer
Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su, and Hui Xiong · 2022
Later among the works it cites.
Maple: Multi-modal prompt learning
Muhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan, and Fahad Shahbaz Khan · 2023
Later among the works it cites.
Self-regulating prompts: Foundational model adaptation without forgetting
Muhammad Uzair Khattak, Syed Talal Wasim, Muzammal Naseer, Salman Khan, Ming-Hsuan Yang, and Fahad Shahbaz Khan · 2023
Later among the works it cites.
Prompt-aligned gradient for prompt tuning
Beier Zhu, Yulei Niu, Yucheng Han, Yue Wu, and Hanwang Zhang · 2023
Later among the works it cites.
Apollo : Unified adapter and prompt learning for vision language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, and Pengchuan Zhang · 2021
Cited alongside, same era.
Towards robust classification model by counterfactual and invariant data generation
Chun-Hao Chang, George-Alexandru Adam, and Anna Goldenberg · 2021
Cited alongside, same era.
Counterfactual zero-shot and open-set visual recognition
Zhongqi Yue, Tan Wang, Qianru Sun, Xian-Sheng Hua, and Hanwang Zhang · 2021
Cited alongside, same era.
Counterfactual generative networks
Axel Sauer and Andreas Geiger · 2021
Cited alongside, same era.
Trusted multi-view classification
Zongbo Han, Changqing Zhang, Huazhu Fu, and Joey Tianyi Zhou · 2021
Cited alongside, same era.
Natural adversarial examples
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song · 2021
Cited alongside, same era.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer · 2021
Cited alongside, same era.
Sanjoy Chowdhury, Sayan Nag, and Dinesh Manocha · 2023
Later among the works it cites.
Argue: Attribute-guided prompt tuning for vision-language models
Xinyu Tian, Shu Zou, Zhaoyuan Yang, and Jing Zhang · 2023
Later among the works it cites.
Counterfactual samples synthesizing and training for robust visual question answering
Long Chen, Yuhang Zheng, Yulei Niu, Hanwang Zhang, and Jun Xiao · 2023
Later among the works it cites.
Disentangle and remerge: Interventional knowledge distillation for few-shot object detection from A conditional causal perspective
Jiangmeng Li, Yanan Zhang, Wenwen Qiang, Lingyu Si, Chengbo Jiao, Xiaohui Hu, Changwen Zheng, and Fuchun Sun · 2023
Later among the works it cites.
Causal balancing for domain generalization
Xinyi Wang, Michael Saxon, Jiachen Li, Hongyang Zhang, Kun Zhang, and William Yang Wang · 2023
Later among the works it cites.
Trusted multi-view classification with dynamic evidential fusion
Zongbo Han, Changqing Zhang, Huazhu Fu, and Joey Tianyi Zhou · 2023
Later among the works it cites.
Consistency-guided prompt learning for vision-language models
Shuvendu Roy and Ali Etemad · 2024
Closest in time.
Clip-adapter: Better vision-language models with feature adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao · 2024
Closest in time.
Sgva-clip: Semantic-guided visual adapting of vision-language models for few-shot image classification
Fang Peng, Xiaoshan Yang, Linhui Xiao, Yaowei Wang, and Changsheng Xu · 2024
Closest in time.
Learning hierarchical prompt with structured linguistic knowledge for vision-language models
Yubin Wang, Xinyang Jiang, De Cheng, Dongsheng Li, and Cairong Zhao · 2024
Closest in time.