Fetching the paper…

VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models · Around