Fetching the paper…
Reading the bibliography…
Chart understanding enables automated data analysis for humans, which requires models to achieve highly accurate visual comprehension.
DVQA: Understanding Data Visualizations via Question Answering
Kafle, K.; Price, B. L.; Cohen, S.; and Kanan, C. 2018 · 2018
Earlier work this paper cites.
FigureQA: An Annotated Figure Dataset for Visual Reasoning
Kahou, S. E.; Michalski, V.; Atkinson, A.; Kádár, Á.; Trischler, A.; and Bengio, Y. 2018 · 2018
Earlier work this paper cites.
ECharts: A declarative framework for rapid construction of web-based visualization
Li, D.; Mei, H.; Shen, Y.; Su, S.; Zhang, W.; Wang, J.; Zu, M.; and Chen, W. 2018 · 2018
Earlier work this paper cites.
Answering Questions about Data Visualizations using Efficient Bimodal Fusion
Kafle, K.; Shrestha, R.; Price, B. L.; Cohen, S.; and Kanan, C. 2020 · 2020
Earlier work this paper cites.
PlotQA: Reasoning over Scientific Plots
Methani, N.; Ganguly, P.; Khapra, M. M.; and Kumar, P. 2020 · 2020
Earlier work this paper cites.
GPT3Mix: Leveraging Large-scale Language Models for Text Augmentation
Yoo, K. M.; Park, D.; Kang, J.; Lee, S.; and Park, W. 2021 · 2021
Earlier work this paper cites.
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Masry, A.; Long, D. X.; Tan, J. Q.; Joty, S. R.; and Hoque, E. 2022 · 2022
Earlier work this paper cites.
RealCQA: Scientific Chart Question Answering as a Test-Bed for First-Order Logic
Ahmed, S.; Jawade, B.; Pandey, S.; Setlur, S.; and Govindaraju, V. 2023 · 2023
Earlier work this paper cites.
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Earlier work this paper cites.
ChartReader: A Unified Framework for Chart Derendering and Comprehension without Heuristic Rules
Cheng, Z.; Dai, Q.; and Hauptmann, A. G. 2023 · 2023
Earlier work this paper cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. C. H. 2023 · 2023
Earlier work this paper cites.
Reinforced Self-Training (ReST) for Language Modeling
Gülçehre, Ç.; Paine, T. L.; Srinivasan, S.; Konyushkova, K.; Weerts, L.; Sharma, A.; Siddhant, A.; Ahern, A.; Wang, M.; Gu, C.; Macherey, W.; Doucet, A.; Firat, O.; and de Freitas, N. 2023 · 2023
Cited alongside, same era.
ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Han, Y.; Zhang, C.; Chen, X.; Yang, X.; Wang, Z.; Yu, G.; Fu, B.; and Zhang, H. 2023 · 2023
Cited alongside, same era.
Lin, Z.; Liu, C.; Zhang, R.; Gao, P.; Qiu, L.; Xiao, H.; Qiu, H.; Lin, C.; Shao, W.; Chen, K.; Han, J.; Huang, S.; Zhang, Y.; He, X.; Li, H.; and Qiao, Y. 2023 · 2023
Cited alongside, same era.
UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning
Masry, A.; Kavehzadeh, P.; Long, D. X.; Hoque, E.; and Joty, S. 2023 · 2023
Cited alongside, same era.
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
Meng, F.; Shao, W.; Lu, Q.; Gao, P.; Zhang, K.; Qiao, Y.; and Luo, P. 2024 · 2024
Closest in time.
OpenAI; Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; and et al., F. L. A. 2024 · 2024
Closest in time.
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
Rose, D.; Himakunthala, V.; Ouyang, A.; He, R.; Mei, A.; Lu, Y.; Saxon, M.; Sonar, C.; Mirza, D.; and Wang, W. Y. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G.; Georgiev, P.; Lei, V. I.; Burnell, R.; Bai, L.; Gulati, A.; Tanzer, G.; Vincent, D.; Pan, Z.; Wang, S.; and et al., S. M. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, W.; Lv, Q.; Yu, W.; Hong, W.; Qi, J.; Wang, Y.; Ji, J.; Yang, Z.; Zhao, L.; Song, X.; Xu, J.; Xu, B.; Li, J.; Dong, Y.; Ding, M.; and Tang, J. 2023 · 2023
Cited alongside, same era.
Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models
Xu, F.; Wu, Z.; Sun, Q.; Ren, S.; Yuan, F.; Yuan, S.; Lin, Q.; Qiao, Y.; and Liu, J. 2023 · 2023
Cited alongside, same era.
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Abdin, M.; Jacobs, S. A.; Awan, A. A.; Aneja, J.; and et al., A. A. 2024 · 2024
Cited alongside, same era.
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Baechler, G.; Sunkara, S.; Wang, M.; Zubach, F.; Mansoor, H.; Etter, V.; Carbune, V.; Lin, J.; Chen, J.; and Sharma, A. 2024 · 2024
Cited alongside, same era.
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs
Carbune, V.; Mansoor, H.; Liu, F.; Aralikatte, R.; Baechler, G.; Chen, J.; and Sharma, A. 2024 · 2024
Cited alongside, same era.
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Li, F.; Zhang, R.; Zhang, H.; Zhang, Y.; Li, B.; Li, W.; Ma, Z.; and Li, C. 2024 · 2024
Cited alongside, same era.
Improved baselines with visual instruction tuning
Liu, H.; Li, C.; Li, Y.; and Lee, Y. J. 2024 · 2024
Cited alongside, same era.
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Chen, Z.; Wu, J.; Wang, W.; Su, W.; Chen, G.; Xing, S.; Zhong, M.; Zhang, Q.; Zhu, X.; Lu, L.; Li, B.; Luo, P.; Lu, T.; Qiao, Y.; and Dai, J. 2023a
Cited in the paper.
Ulmer, D.; Mansimov, E.; Lin, K.; Sun, J.; Gao, X.; and Zhang, Y. 2024 · 2024
Closest in time.
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
Wang, Z.; Xia, M.; He, L.; Chen, H.; Liu, Y.; Zhu, R.; Liang, K.; Wu, X.; Liu, H.; Malladi, S.; Chevalier, A.; Arora, S.; and Chen, D. 2024 · 2024
Closest in time.
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
Xia, R.; Zhang, B.; Ye, H.; Yan, X.; Liu, Q.; Zhou, H.; Chen, Z.; Dou, M.; Shi, B.; Yan, J.; and Qiao, Y. 2024 · 2024
Closest in time.
Interactive Evolution: A Neural-Symbolic Self-Training Framework For Large Language Models
Xu, F.; Sun, Q.; Cheng, K.; Liu, J.; Qiao, Y.; and Wu, Z. 2024 · 2024
Closest in time.
TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning
Zhang, L.; Hu, A.; Xu, H.; Yan, M.; Xu, Y.; Jin, Q.; Zhang, J.; and Huang, F. 2024 · 2024
Closest in time.
Efficient End-to-End Visual Document Understanding with Rationale Distillation
Zhu, W.; Agarwal, A.; Joshi, M.; Jia, R.; Thomason, J.; and Toutanova, K. 2024 · 2024
Closest in time.