Fetching the paper…
Reading the bibliography…
In the fields of computer vision and natural language processing, multimodal chart question-answering, especially involving color, structure, and textless charts, poses significant challenges.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., Berg, T., 2011 · 2011
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes, in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp. 787–798
Kazemzadeh, S., Ordonez, V., Matten, M., Berg, T., 2014 · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.Y., Vedantam, R., Gupta, S., Dollár, P., Zitnick, C.L., 2015 · 2015
Earlier work this paper cites.
Datatone: Managing ambiguity in natural language interfaces for data visualization, in: Proceedings of the 28th annual acm symposium on user interface software & technology, pp. 489–500
Gao, T., Dontcheva, M., Adar, E., Liu, Z., Karahalios, K.G., 2015 · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 11–20
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A.L., Murphy, K., 2016 · 2016
Earlier work this paper cites.
Eviza: A natural language interface for visual analysis, in: Proceedings of the 29th annual symposium on user interface software and technology, pp. 365–377
Setlur, V., Battersby, S.E., Tory, M., Gossweiler, R., Chang, A.X., 2016 · 2016
Earlier work this paper cites.
Applying pragmatics principles for interaction with visual analytics
Hoque, E., Setlur, V., Tory, M., Dykeman, I., 2017 · 2017
Earlier work this paper cites.
Figureqa: An annotated figure dataset for visual reasoning
Kahou, S.E., Michalski, V., Atkinson, A., Kádár, Á., Trischler, A., Bengio, Y., 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., et al., 2017 · 2017
Earlier work this paper cites.
Orko: Facilitating multimodal interaction for visual exploration and analysis of networks
Srinivasan, A., Stasko, J., 2017 · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning, in: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2556–2565
Sharma, P., Ding, N., Goodman, S., Soricut, R., 2018 · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale, in: Proceedings of the IEEE/CVF international conference on computer vision, pp. 8948–8957
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., Anderson, P., 2019 · 2019
Earlier work this paper cites.
Figurenet: A deep learning model for question-answering on scientific plots, in: 2019 International Joint Conference on Neural Networks (IJCNN), IEEE. pp. 1–8
Reddy, R., Ramesh, R., Deshpande, A., Khapra, M.M., 2019 · 2019
Cited alongside, same era.
Leaf-qa: Locate, encode & attend for figure question answering, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 3512–3521
Chaudhry, R., Shekhar, S., Gupta, U., Maneriker, P., Bansal, P., Joshi, A., 2020 · 2020
Cited alongside, same era.
Answering questions about data visualizations using efficient bimodal fusion, in: Proceedings of the IEEE/CVF Winter conference on applications of computer vision, pp. 1498–1507
Kafle, K., Shrestha, R., Cohen, S., Price, B., Kanan, C., 2020 · 2020
Cited alongside, same era.
Plotqa: Reasoning over scientific plots, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 1527–1536
Methani, N., Ganguly, P., Khapra, M.M., Kumar, P., 2020 · 2020
Cited alongside, same era.
A novel extended multimodal ai framework towards vulnerability detection in smart contracts
Jie, W., Chen, Q., Wang, J., Koe, A.S.V., Li, J., Huang, P., Wu, Y., Wang, Y., 2023 · 2023
Later among the works it cites.
Pix2struct: Screenshot parsing as pretraining for visual language understanding, in: International Conference on Machine Learning, PMLR. pp. 18893–18912
Lee, K., Joshi, M., Turc, I.R., Hu, H., Liu, F., Eisenschlos, J.M., Khandelwal, U., Shaw, P., Chang, M.W., Toutanova, K., 2023 · 2023
Later among the works it cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, in: International conference on machine learning, PMLR. pp. 19730–19742
Li, J., Li, D., Savarese, S., Hoi, S., 2023 · 2023
Later among the works it cites.
DePlot: One-shot visual language reasoning by plot-to-table translation, in: Rogers, A., Boyd-Graber, J., Okazaki, N. (Eds.), Findings of the Association for Computational Linguistics: ACL 2023, Association for Computational Linguistics, Toronto, Canada. pp. 10381–10399
Liu, F., Eisenschlos, J., Piccinno, F., Krichene, S., Pang, C., Lee, K., Joshi, M., Chen, W., Collier, N., Altun, Y., 2023a · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stl-cqa: Structure-based transformers with localization and encoding for chart question answering, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 3275–3284
Singh, H., Shekhar, S., 2020 · 2020
Cited alongside, same era.
Mufasa: Multimodal fusion architecture search for electronic health records, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 10532–10540
Xu, Z., So, D.R., Dai, A.M., 2021 · 2021
Cited alongside, same era.
Chart question answering: State of the art and future directions, in: Computer Graphics Forum, Wiley Online Library. pp. 555–572
Hoque, E., Kavehzadeh, P., Masry, A., 2022 · 2022
Cited alongside, same era.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning, in: Muresan, S., Nakov, P., Villavicencio, A. (Eds.), Findings of the Association for Computational Linguistics: ACL 2022, Association for Computational Linguistics, Dublin, Ireland. pp. 2263–2279
Masry, A., Do, X.L., Tan, J.Q., Joty, S., Hoque, E., 2022 · 2022
Cited alongside, same era.
A multimodal computer-aided diagnostic system for depression relapse prediction using audiovisual cues: A proof of concept
Othmani, A., Zeghina, A.O., 2022 · 2022
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., Zhou, J., 2023 · 2023
Cited alongside, same era.
Chartreader: A unified framework for chart derendering and comprehension without heuristic rules, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 22202–22213
Cheng, Z.Q., Dai, Q., Hauptmann, A.G., 2023 · 2023
Cited alongside, same era.
Multimodal document analytics for banking process automation
Gerling, C., Lessmann, S., 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Unichart: A universal vision-language pretrained model for chart comprehension and reasoning
Masry, A., Kavehzadeh, P., Do, X.L., Hoque, E., Joty, S., 2023 · 2023
Later among the works it cites.
The impact of multimodal large language models on health care’s future
Meskó, B., 2023 · 2023
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world
Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., Wei, F., 2023 · 2023
Later among the works it cites.
Zhou, M., Fung, Y., Chen, L., Thomas, C., Ji, H., Chang, S.F., 2023 · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A.M.H., Zhao, J., Wang, W., Li, B., Fung, P.N., Hoi, S., 2024 · 2024
Closest in time.
Meng, F., Shao, W., Lu, Q., Gao, P., Zhang, K., Qiao, Y., Luo, P., 2024 · 2024
Closest in time.
Fusion of electronic health records and radiographic images for a multimodal deep learning prediction model of atypical femur fractures
Schilcher, J., Nilsson, A., Andlid, O., Eklund, A., 2024 · 2024
Closest in time.
Chartx & chartvlm: A versatile benchmark and foundation model for complicated chart reasoning
Xia, R., Zhang, B., Ye, H., Yan, X., Liu, Q., Zhou, H., Chen, Z., Dou, M., Shi, B., Yan, J., et al., 2024 · 2024
Closest in time.