Fetching the paper…
Reading the bibliography…
Charts are common in literature across various scientific fields, conveying rich information easily accessible to readers.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in the 34th International Conference on Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 1877–1901
1901
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of Machine Learning Research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
M. Savva, N. Kong, A. Chhajta, L. Fei-Fei, M. Agrawala, and J. Heer, “Revision: automated classification, analysis and redesign of chart images,” Proceedings of the 24th annual ACM symposium on User interface software and technology , 2011
2011
Earlier work this paper cites.
P. Pasupat and P. Liang, “Compositional semantic parsing on semi-structured tables,” in the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing , 2015
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Law and J. Deng, “Cornernet: Detecting objects as paired keypoints,” International Journal of Computer Vision , vol. 128, pp. 642 – 656, 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Choi, S. Jung, D. G. Park, J. Choo, and N. Elmqvist, “Visualizing for the non-visual: Enabling the visually impaired to use visualization,” Computer Graphics Forum , vol. 38, 2019
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in Advances in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar, “Plotqa: Reasoning over scientific plots,” 2020 IEEE Winter Conference on Applications of Computer Vision (WACV) , pp. 1516–1525, 2019
2019
Earlier work this paper cites.
J. Obeid and E. Hoque, “Chart-to-text: Generating natural language descriptions for charts by adapting the transformer model,” in International Conference on Natural Language Generation , 2020
2020
Earlier work this paper cites.
D. Ulmer, L. Meijerink, and G. Cinà, “Trust issues: Uncertainty estimation does not enable reliable ood detection on medical tabular data,” in Machine Learning for Health . PMLR, 2020, pp. 341–354
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai, “Vl-bert: Pre-training of generic visual-linguistic representations,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
J. Herzig, P. K. Nowak, T. Müller, F. Piccinno, and J. M. Eisenschlos, “Tapas: Weakly supervised table parsing via pre-training,” in Annual Meeting of the Association for Computational Linguistics , 2020
2020
Earlier work this paper cites.
J. Luo, Z. Li, J. Wang, and C.-Y. Lin, “Chartocr: Data extraction from charts images via a deep hybrid framework,” 2021 IEEE Winter Conference on Applications of Computer Vision (WACV) , pp. 1916–1924, 2021
2021
Earlier work this paper cites.
C. Rane, S. M. Subramanya, D. S. Endluri, J. Wu, and C. L. Giles, “Chartreader: Automatic parsing of bar-plots,” 2021 IEEE 22nd International Conference on Information Reuse and Integration for Data Science (IRI) , pp. 318–325, 2021
2021
Earlier work this paper cites.
W. Li, C. Gao, G. Niu, X. Xiao, H. Liu, J. Liu, H. Wu, and H. Wang, “Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning,” the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing , 2021
2021
Earlier work this paper cites.
Z. Huang, Z. Zeng, Y. Huang, B. Liu, D. Fu, and J. Fu, “Seeing out of the box: End-to-end pre-training for vision-language representation learning,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 12 971–12 980, 2021
2021
Earlier work this paper cites.
J. Cho, J. Lei, H. Tan, and M. Bansal, “Unifying vision-and-language tasks via text generation,” in International Conference on Machine Learning , 2021
2021
Earlier work this paper cites.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. V. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in International Conference on Machine Learning , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
W. Kim, B. Son, and I. Kim, “Vilt: Vision-and-language transformer without convolution or region supervision,” in International Conference on Machine Learning , 2021
2021
Cited alongside, same era.
H. Xue, Y. Huang, B. Liu, H. Peng, J. Fu, H. Li, and J. Luo, “Probing inter-modality: Visual parsing with self-attention for vision-language pre-training,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
OpenAI, “Gpt-4v,” https://openai.com/index/gpt-4v-system-card/ , 2023
2023
Closest in time.
B. Zhang, J. Yuan, B. Shi, T. Chen, Y. Li, and Y. Qiao, “Uni3d: A unified baseline for multi-dataset 3d object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 9253–9262
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Li, R. R. Selvaraju, A. D. Gotmare, S. R. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela, “Flava: A foundational language and vision alignment model,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 15 617–15 629, 2021
2021
Cited alongside, same era.
L. Nan, C.-H. Hsieh, Z. Mao, X. V. Lin, N. Verma, R. Zhang, W. Kryscinski, N. Schoelkopf, R. Kong, X. Tang, M. Mutuma, B. Rosand, I. Trindade, R. Bandaru, J. Cunningham, C. Xiong, and D. R. Radev, “Fetaqa: Free-form table question answering,” Transactions of the Association for Computational Linguistics , vol. 10, pp. 35–49, 2021
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
A. Masry, X. L. Do, J. Q. Tan, S. R. Joty, and E. Hoque, “Chartqa: A benchmark for question answering about charts with visual and logical reasoning,” in Findings of the Association for Computational Linguistics: ACL , 2022
2022
Cited alongside, same era.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 16 000–16 009
2022
Cited alongside, same era.
F. Liu, F. Piccinno, S. Krichene, C. Pang, K. Lee, M. Joshi, Y. Altun, N. Collier, and J. M. Eisenschlos, “Matcha: Enhancing visual language pretraining with math reasoning and chart derendering,” in Annual Meeting of the Association for Computational Linguistics , 2022
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
M. Zhou, Y. R. Fung, L. Chen, C. Thomas, H. Ji, and S.-F. Chang, “Enhanced chart understanding in vision and language task via cross-modal pre-training on plot table pairs,” in the Association for Computational Linguistics: ACL , 2023
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” ArXiv , vol. abs/2303.08774, 2023
2023
Closest in time.
X. Chen, X. Wang, S. Changpinyo, A. J. Piergiovanni, P. Padlewski, D. M. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, A. Kolesnikov, J. Puigcerver, N. Ding, K. Rong, H. Akbari, G. Mishra, L. Xue, A. V. Thapliyal, J. Bradbury, W. Kuo, M. Seyedhosseini, C. Jia, B. K. Ayan, C. Riquelme, A. Steiner, A. Angelova, X. Zhai, N. Houlsby, and R. Soricut, “Pali: A jointly-scaled multilingual language-image model,” in International Conference on Learning Representations , 2023
2023
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Advances in Neural Information Processing Systems , 2023
2023
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma et al. , “Scaling instruction-finetuned language models,” The Journal of Machine Learning Research , 2024
2024
Closest in time.
R. Xia, B. Zhang, H. Ye, X. Yan, Q. Liu, H. Zhou, Z. Chen, M. Dou, B. Shi, J. Yan et al. , “Chartx & chartvlm: A versatile benchmark and foundation model for complicated chart reasoning,” ArXiv , vol. 2402.12185, 2024
2024
Closest in time.
R. Lou, K. Zhang, J. Xie, Y. Sun, J. Ahn, H. Xu, Y. Su, and W. Yin, “Muffin: Curating multi-faceted instructions for improving instruction-following,” in International Conference on Learning Representations , 2024
2024
Closest in time.
Y. Li, C. Zhang, G. Yu, Z. Wang, B. Fu, G. Lin, C. Shen, L. Chen, and Y. Wei, “Stablellava: Enhanced visual instruction tuning with synthesized image-dialogue data,” in Findings of the Association for Computational Linguistics: ACL , 2024
2024
Closest in time.
J. Zhan, J. Dai, J. Ye, Y. Zhou, D. Zhang, Z. Liu, X. Zhang, R. Yuan, G. Zhang, L. Li, H. Yan, J. Fu, T. Gui, T. Sun, Y. Jiang, and X. Qiu, “Anygpt: Unified multimodal llm with discrete sequence modeling,” in the 62nd Annual Meeting of the Association for Computational Linguistics , 2024
2024
Closest in time.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
Closest in time.