Fetching the paper…
Reading the bibliography…
Interpretability in Vision-Language Models (VLMs) is crucial for trust, debugging, and decision-making in high-stakes applications.
A value for n-person games
Lloyd S. Shapley · 1953
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
Why should i trust you? explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
R. Fong and A. Vedaldi · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Rise: Randomized input sampling for explanation of black-box models
V. Petsiuk, A. Das, and K. Saenko · 2018
Earlier work this paper cites.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Earlier work this paper cites.
Understanding deep networks via extremal perturbations and smooth masks
R. Fong and A. Vedaldi · 2019
Earlier work this paper cites.
Xrai: Better attributions through regions
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2019
Cited alongside, same era.
Xatten: Interpretable cross-modal attention analysis for vision and language reasoning
T. Gokhale, R. Goyal, C. Baral, and M. Bansal · 2020
Cited alongside, same era.
Score-cam: Score-weighted visual explanations for convolutional neural networks
Haofan Wang, Zifan Wang, Mengnan Du, Fan Yang, Zijian Zhang, Sirui Ding, Piotr Mardziel, and Xia Hu · 2020
Cited alongside, same era.
Evaluating bias in vision-language models with synthetic counterfactuals
S. Agarwal, I. Malkiel, M. Donini, L. Zhou, J. K. Kummerfeld, R. Mihalcea, and L. Zitnick · 2021
Cited alongside, same era.
Transformer interpretability beyond attention visualization
Hila Chefer, Shir Gur, and Lior Wolf · 2021
Cited alongside, same era.
Fastshap: Real-time shapley value estimation
Harsha Jethani, Mukund Sundararajan, Sumit Basu, and John Duchi · 2021
Mm-shap: Multimodal shapley values for model interpretation
L. Parcalabescu and A. Frank · 2022
Later among the works it cites.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Yuxin Rolland, Linus Gustafson, Trevor Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick · 2023
Later among the works it cites.
Gemini 2.0 pro, 2024
Google DeepMind · 2024
Later among the works it cites.
Tokenshap: Interpreting large language models with monte carlo shapley value estimation
Miriam Horovicz and Roni Goldshmidt · 2024
Later among the works it cites.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2024
Later among the works it cites.
Gpt-4o model card, 2024
OpenAI · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D-rise: Dynamic randomized input sampling for object detection explanations
V. Petsiuk, A. Das, and K. Saenko · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pam Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Feng Li, Hao Zhang, Xiao Zhang, Lei Zhu, Hang Wang, Jianlong Shi, Hongyang Li, and Hao Dong
Cited in the paper.
Llama 3.2: An open-weight vision-language model
Meta AI Research · 2024
Later among the works it cites.