Fetching the paper…
Reading the bibliography…
Building a large-scale figure QA dataset requires a considerable amount of work, from gathering and selecting figures to extracting attributes like text, numbers, and colors, and generating QAs.
Unifying Vision-and-Language Tasks via Text Generation
Cho, J.; Lei, J.; Tan, H.; and Bansal, M. 2021 · 1942
Earlier work this paper cites.
FigureSeer: Parsing Result-Figures in Research Papers
Siegel, N.; Horvitz, Z.; Levin, R.; Divvala, S. K.; and Farhadi, A. 2016 · 2016
Earlier work this paper cites.
DVQA: Understanding Data Visualizations via Question Answering
Kafle, K.; Price, B.; Cohen, S.; and Kanan, C. 2018 · 2018
Earlier work this paper cites.
FigureQA: An Annotated Figure Dataset for Visual Reasoning
Kahou, S. E.; Michalski, V.; Atkinson, A.; Kadar, A.; Trischler, A.; and Bengio, Y. 2018 · 2018
Earlier work this paper cites.
LEAF-QA: Locate, Encode & Attend for Figure Question Answering
Chaudhry, R.; Shekhar, S.; Gupta, U.; Maneriker, P.; Bansal, P.; and Joshi, A. 2020 · 2020
Earlier work this paper cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2020 · 2020
Earlier work this paper cites.
PlotQA: Reasoning over Scientific Plots
Methani, N.; Ganguly, P.; Khapra, M. M.; and Kumar, P. 2020 · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Cited alongside, same era.
STL-CQA: Structure-based Transformers with Localization and Encoding for Chart Question Answering
Singh, H.; and Shekhar, S. 2020 · 2020
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Cited alongside, same era.
Chart question answering: State of the art and future directions
Hoque, E.; Kavehzadeh, P.; and Masry, A. 2022 · 2022
Cited alongside, same era.
ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Han, Y.; Zhang, C.; Chen, X.; Yang, X.; Wang, Z.; Yu, G.; Fu, B.; and Zhang, H. 2023 · 2023
Later among the works it cites.
Pix2Struct: screenshot parsing as pretraining for visual language understanding
Lee, K.; Joshi, M.; Turc, I.; Hu, H.; Liu, F.; Eisenschlos, J.; Khandelwal, U.; Shaw, P.; Chang, M.-W.; and Toutanova, K. 2023 · 2023
Later among the works it cites.
UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning
Masry, A.; Kavehzadeh, P.; Do, X. L.; Hoque, E.; and Joty, S. 2023 · 2023
Later among the works it cites.
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs
Carbune, V.; Mansoor, H.; Liu, F.; Aralikatte, R.; Baechler, G.; Chen, J.; and Sharma, A. 2024 · 2024
Closest in time.
Synthesize Step-by-Step: Tools Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
OCR-Free Document Understanding Transformer
Kim, G.; Hong, T.; Yim, M.; Nam, J.; Park, J.; Yim, J.; Hwang, W.; Yun, S.; Han, D.; and Park, S. 2022 · 2022
Cited alongside, same era.
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Masry, A.; Long, D.; Tan, J. Q.; Joty, S.; and Hoque, E. 2022 · 2022
Cited alongside, same era.
DePlot: One-shot visual language reasoning by plot-to-table translation
Liu, F.; Eisenschlos, J.; Piccinno, F.; Krichene, S.; Pang, C.; Lee, K.; Joshi, M.; Chen, W.; Collier, N.; and Altun, Y. 2023a
Cited in the paper.
MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering
Liu, F.; Piccinno, F.; Krichene, S.; Pang, C.; Lee, K.; Joshi, M.; Altun, Y.; Collier, N.; and Eisenschlos, J. 2023b
Cited in the paper.
Li, Z.; Jasani, B.; Tang, P.; and Ghadar, S. 2024 · 2024
Closest in time.
Gemma: Open Models Based on Gemini Research and Technology
Team, G.; Mesnard, T.; Hardin, C.; Dadashi, R.; Bhupatiraju, S.; Pathak, S.; Sifre, L.; Rivière, M.; Kale, M. S.; Love, J.; Tafti, P.; Hussenot, L.; Sessa, P. G.; Chowdhery, A.; Roberts, A.; Barua, A.; Botev, A.; Castro-Ros, A.; Slone, A.; Héliou, A.; Tacchetti, A.; Bulanova, A.; Paterson, A.; Tsai, B.; Shahriari, B.; Lan, C. L.; Choquette-Choo, C. A.; Crepy, C.; Cer, D.; Ippolito, D.; Reid, D.; Buchatskaya, E.; Ni, E.; Noland, E.; Yan, G.; Tucker, G.; Muraru, G.-C.; Rozhdestvenskiy, G.; Michalewski, H.; Tenney, I.; Grishchenko, I.; Austin, J.; Keeling, J.; Labanowski, J.; Lespiau, J.-B.; Stanway, J.; Brennan, J.; Chen, J.; Ferret, J.; Chiu, J.; Mao-Jones, J.; Lee, K.; Yu, K.; Millican, K.; Sjoesund, L. L.; Lee, L.; Dixon, L.; Reid, M.; Mikuła, M.; Wirth, M.; Sharman, M.; Chinaev, N.; Thain, N.; Bachem, O.; Chang, O.; Wahltinez, O.; Bailey, P.; Michel, P.; Yotov, P.; Chaabouni, R.; Comanescu, R.; Jana, R.; Anil, R.; McIlroy, R.; Liu, R.; Mullins, R.; Smith, S. L.; Borgeaud, S.; Girgin, S.; Douglas, S.; Pandya, S.; Shakeri, S.; De, S.; Klimenko, T.; Hennigan, T.; Feinberg, V.; Stokowiec, W.; hui Chen, Y.; Ahmed, Z.; Gong, Z.; Warkentin, T.; Peran, L.; Giang, M.; Farabet, C.; Vinyals, O.; Dean, J.; Kavukcuoglu, K.; Hassabis, D.; Ghahramani, Z.; Eck, D.; Barral, J.; Pereira, F.; Collins, E.; Joulin, A.; Fiedel, N.; Senter, E.; Andreev, A.; and Kenealy, K. 2024 · 2024
Closest in time.