Fetching the paper…
Reading the bibliography…
Multimodal pre-trained models, such as CLIP, are popular for zero-shot classification due to their open-vocabulary flexibility and high performance.
MIT press, 1992
J. H. Holland, Adaptation in natural and artificial systems: an introductory analysis with applications to biology, control, and artificial intelligence · 1992
Earlier work this paper cites.
V. S. Ramachandran and E. M. Hubbard, “Synaesthesia–a window into perception, thought and language,” Journal of consciousness studies
2001
Earlier work this paper cites.
P. Melville and R. J. Mooney, “Diverse ensembles for active learning,” in Proceedings of the twenty-first international conference on Machine learning
2004
Earlier work this paper cites.
V. Ferrari and A. Zisserman, “Learning visual attributes,” Advances in neural information processing systems
2007
Earlier work this paper cites.
A. Farhadi, I. Endres, D. Hoiem, and D. Forsyth, “Describing objects by their attributes,” in 2009 IEEE conference on computer vision and pattern recognition
2009
Earlier work this paper cites.
C. H. Lampert, H. Nickisch, and S. Harmeling, “Attribute-based classification for zero-shot visual object categorization,” 2013
2013
Earlier work this paper cites.
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov, “Devise: A deep visual-semantic embedding model,” in NeurIPS
2013
Earlier work this paper cites.
K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep inside convolutional networks: Visualising image classification models and saliency maps,” CoRR
2013
Earlier work this paper cites.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in ECCV
2014
Earlier work this paper cites.
Z. Akata, S. Reed, D. Walter, H. Lee, and B. Schiele, “Evaluation of output embeddings for fine-grained image classification,” in CVPR
2015
Earlier work this paper cites.
B. Romera-Paredes and P. Torr, “An embarrassingly simple approach to zero-shot learning,” in ICML
2015
Earlier work this paper cites.
S. Huang, Z. Xu, D. Tao, and Y. Zhang, “Part-stacked cnn for fine-grained visual categorization,” in CVPR
2016
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in CVPR
2016
Earlier work this paper cites.
A. M. Nguyen, A. Dosovitskiy, J. Yosinski, T. Brox, and J. Clune, “Synthesizing the preferred inputs for neurons in neural networks via deep generator networks,” in NeurIPS
2016
Earlier work this paper cites.
E. Kodirov, T. Xiang, and S. Gong, “Semantic autoencoder for zero-shot learning,” in CVPR
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in ICCV
2017
Earlier work this paper cites.
P. W. Koh and P. Liang, “Understanding black-box predictions via influence functions,” in ICML
2017
Earlier work this paper cites.
B. Zhou, Y. Sun, D. Bau, and A. Torralba, “Interpretable basis decomposition for visual explanation,” in ECCV
2018
Earlier work this paper cites.
V. Petsiuk, A. Das, and K. Saenko, “Rise: Randomized input sampling for explanation of black-box models,” CoRR
2018
Earlier work this paper cites.
C.-K. Yeh, J. Kim, I. E.-H. Yen, and P. K. Ravikumar, “Representer point selection for explaining deep neural networks,” NeurIPS
2018
Earlier work this paper cites.
J. Frankle and M. Carbin, “The lottery ticket hypothesis: Training pruned neural networks,” CoRR
2018
Earlier work this paper cites.
R. Fong, M. Patrick, and A. Vedaldi, “Understanding deep networks via extremal perturbations and smooth masks,” in ICCV
2019
Earlier work this paper cites.
Y. Goyal, Z. Wu, J. Ernst, D. Batra, D. Parikh, and S. Lee, “Counterfactual visual explanations,” in ICML
2019
Earlier work this paper cites.
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang, “Visualbert: A simple and performant baseline for vision and language,” CoRR
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” CoRR
2019
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” NeurIPS
2019
Cited alongside, same era.
P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang, “Concept bottleneck models,” in ICML
2020
Cited alongside, same era.
L. Tang, D. Wertheimer, and B. Hariharan, “Revisiting pose-normalization for fine-grained few-shot recognition,” in CVPR
2020
Cited alongside, same era.
P. Wang and N. Vasconcelos, “Scout: Self-aware discriminant counterfactual explanations,” in CVPR
2020
Cited alongside, same era.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al
2022
Later among the works it cites.
J. Chen, H. Guo, K. Yi, B. Li, and M. Elhoseiny, “Visualgpt: Data-efficient adaptation of pretrained language models for image captioning,” in CVPR
2022
Later among the works it cites.
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, et al
2022
Later among the works it cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICML
2022
Later among the works it cites.
Z. Luo, Y. Xi, R. Zhang, and J. Ma, “A frustratingly simple approach for end-to-end image captioning,” CoRR
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Pruthi, F. Liu, M. Sundararajan, and S. Kale, “Estimating training data influence by tracking gradient descent,” CoRR
2020
Cited alongside, same era.
A. Silva, R. Chopra, and M. C. Gombolay, “Cross-loss influence functions to explain deep network representations,” in AISTATS
2020
Cited alongside, same era.
H. Guo, N. Rajani, P. Hase, M. Bansal, and C. Xiong, “Fastif: Scalable influence functions for efficient model interpretation and debugging,” CoRR
2020
Cited alongside, same era.
B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, et al
2020
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al
2021
Cited alongside, same era.
G. Van Horn, E. Cole, S. Beery, K. Wilber, S. Belongie, and O. Mac Aodha, “Benchmarking representation learning for natural world image collections,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
2021
Cited alongside, same era.
H. Chefer, S. Gur, and L. Wolf, “Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers,” in ICCV
2021
Cited alongside, same era.
S. Menon and C. Vondrick, “Visual classification via description from large language models,” ICLR
2023
Later among the works it cites.
K. Roth, J. M. Kim, A. Koepke, O. Vinyals, C. Schmid, and Z. Akata, “Waffling around for performance: Visual classification with random words and broad concepts,” CoRR
2023
Later among the works it cites.
C.-P. Tsai, C.-K. Yeh, and P. Ravikumar, “Sample based explanations via generalized representers,” CoRR
2023
Later among the works it cites.
Y. Gandelsman, A. A. Efros, and J. Steinhardt, “Interpreting clip’s image representation via text-based decomposition,” 2023
2023
Later among the works it cites.
A. Dravid, Y. Gandelsman, A. A. Efros, and A. Shocher, “Rosetta neurons: Mining the common units in a model zoo,” in CVPR
2023
Later among the works it cites.
M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev, “Reproducible scaling laws for contrastive language-image learning,” in CVPR
2023
Later among the works it cites.
X. Xu, C. Wu, S. Rosenman, V. Lal, W. Che, and N. Duan, “Bridgetower: Building bridges between encoders in vision-language representation learning,” in AAAI
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” CoRR
2023
Later among the works it cites.
S. Pratt, I. Covert, R. Liu, and A. Farhadi, “What does a platypus look like? generating customized prompts for zero-shot image classification,” in Proceedings of the IEEE/CVF International Conference on Computer Vision
2023
Later among the works it cites.
A. Yan, Y. Wang, Y. Zhong, C. Dong, Z. He, Y. Lu, W. Y. Wang, J. Shang, and J. McAuley, “Learning concise and descriptive attributes for visual recognition,” in CVPR
2023
Later among the works it cites.
B. Romera-Paredes, M. Barekatain, A. Novikov, M. Balog, M. P. Kumar, E. Dupont, F. J. Ruiz, J. S. Ellenberg, P. Wang, O. Fawzi, et al
2023
Later among the works it cites.
S. Yu, S. Liu, Z. Lin, D. Pathak, and D. Ramanan, “Language models as black-box optimizers for vision-language models,” CoRR
2023
Later among the works it cites.
S. Han, L. Zhuo, Y. Liao, and S. Liu, “Llms as visual explainers: Advancing image classification with evolving visual descriptions,” CoRR
2023
Later among the works it cites.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al
2023
Later among the works it cites.
M. Alper and H. Averbuch-Elor, “Kiki or bouba? sound symbolism in vision-and-language models,” Advances in Neural Information Processing Systems
2024
Closest in time.
V. Prabhu, S. Yenamandra, P. Chattopadhyay, and J. Hoffman, “Lance: Stress-testing visual models by generating language-guided counterfactual images,” NeurIPS
2024
Closest in time.
L. Li, Z.-Y. Dou, N. Peng, and K.-W. Chang, “Desco: Learning object recognition with rich language descriptions,” NeurIPS
2024
Closest in time.
S. Jin, X. Jiang, J. Huang, L. Lu, and S. Lu, “Llms meet vlms: Boost open vocabulary object detection with fine-grained descriptors,” CoRR
2024
Closest in time.