Fetching the paper…

Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM · Around