Fetching the paper…
Reading the bibliography…
Despite the promising results of large multimodal models (LMMs) in complex vision-language tasks that require knowledge, reasoning, and perception abilities together, we surprisingly found that these models struggle with simple tasks on infographics that require perception only.
The Semiology of Graphics
J. Bertin · 1967
Earlier work this paper cites.
Graphical perception: Theory, experimentation, and application to the development of graphical methods
William S Cleveland and Robert McGill · 1984
Earlier work this paper cites.
Sizing the horizon: the effects of chart size and layering on the graphical perception of time series visualizations
Jeffrey Heer, Nicholas Kong, and Maneesh Agrawala · 2009
Earlier work this paper cites.
Crowdsourcing graphical perception: using mechanical turk to assess visualization design
Jeffrey Heer and Michael Bostock · 2010
Earlier work this paper cites.
Graphical perception of multiple time series
Waqas Javed, Bryan McDonnel, and Niklas Elmqvist · 2010
Earlier work this paper cites.
What makes a visualization memorable?
Michelle A. Borkin, Azalea A. Vo, Zoya Bylinskii, Phillip Isola, Shashank Sunkavalli, Aude Oliva, and Hanspeter Pfister · 2013
Earlier work this paper cites.
Visualization Analysis and Design
Tamara Munzner · 2014
Earlier work this paper cites.
” why should i trust you?” explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott Lundberg · 2017
Earlier work this paper cites.
Vega-lite: A grammar of interactive graphics
Arvind Satyanarayan, Dominik Moritz, Kanit Wongsuphasawat, and Jeffrey Heer · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Evaluating ‘graphical perception’ with cnns
Daniel Haehn, James Tompkin, and Hanspeter Pfister · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Task-based effectiveness of basic visualizations
Bahador Saket, Alex Endert, and Çağatay Demiralp · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Cited alongside, same era.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Graphical perception for immersive analytics
Matt Whitlock, Stephen Smart, and Danielle Albers Szafir · 2020
Cited alongside, same era.
Seeing what you believe or believing what you see? belief biases correlation estimation
Cindy Xiong, Chase Stokes, Yea-Seul Kim, and Steven Franconeri · 2023
Later among the works it cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang · 2023
Later among the works it cites.
What does the chart say? grouping cues guide viewer comparisons and conclusions in bar charts
Cindy Xiong Bearfield, Chase Stokes, Andrew Lovett, and Steven Franconeri · 2024
Later among the works it cites.
Mechanistic interpretability for AI safety - a review
Leonard Bereska and Stratis Gavves · 2024
Later among the works it cites.
Making large multimodal models understand arbitrary visual prompts
Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Do Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque · 2022
Cited alongside, same era.
Infographicvqa
Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and C. V. Jawahar · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2022
Cited alongside, same era.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Google Gemini Team · 2023
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Cited alongside, same era.
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Google Gemini Team · 2024
Later among the works it cites.
Natural language dataset generation framework for visualizations powered by large language models
Hyung-Kwon Ko, Hyeon Jeon, Gwanmo Park, Dae Hyun Kim, Nam Wook Kim, Juho Kim, and Jinwook Seo · 2024
Later among the works it cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2024
Later among the works it cites.
Chartassisstant: A universal chart multimodal language model via chart-to-table pre-training and multitask instruction tuning
Fanqing Meng, Wenqi Shao, Quanfeng Lu, Peng Gao, Kaipeng Zhang, Yu Qiao, and Ping Luo · 2024
Later among the works it cites.
Rethinking interpretability in the era of large language models
Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao · 2024
Later among the works it cites.
Introducing the next generation of claude
Claude Team · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann LeCun, and Saining Xie · 2024
Later among the works it cites.
Magiclens: Self-supervised image retrieval with open-ended instructions
Kai Zhang, Yi Luan, Hexiang Hu, Kenton Lee, Siyuan Qiao, Wenhu Chen, Yu Su, and Ming-Wei Chang · 2024
Later among the works it cites.