Fetching the paper…
Reading the bibliography…
Chart comprehension presents significant challenges for machine learning models due to the diverse and intricate shapes of charts.
PlotQA: Reasoning over Scientific Plots, Feb. 2020
Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra, and Pratyush Kumar · 1909
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, July 2020
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 1910
Earlier work this paper cites.
Multilingual Denoising Pre-training for Neural Machine Translation, Jan. 2020
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer · 2001
Earlier work this paper cites.
SOLOv2: Dynamic and Fast Instance Segmentation, Oct. 2020
Xinlong Wang, Rufeng Zhang, Tao Kong, Lei Li, and Chunhua Shen · 2003
Earlier work this paper cites.
TAPAS: Weakly Supervised Table Parsing via Pre-training
Jonathan Herzig, Paweł Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos · 2004
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, June 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Deep Residual Learning for Image Recognition, Dec. 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Deep Learning in Neural Networks: An Overview
Juergen Schmidhuber · 2015
Earlier work this paper cites.
VQA: Visual Question Answering, Oct. 2016
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Jan. 2016
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2016
Earlier work this paper cites.
Stacked Attention Networks for Image Question Answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2016
Earlier work this paper cites.
Deformable Convolutional Networks, June 2017
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei · 2017
Earlier work this paper cites.
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2018
Earlier work this paper cites.
Visualizing for the Non‐Visual: Enabling the Visually Impaired to Use Visualization
Jinho Choi, Sanghun Jung, Deok Gun Park, Jaegul Choo, and Niklas Elmqvist · 2019
Earlier work this paper cites.
Data extraction from charts via single deep neural network, 2019
Xiaoyi Liu, Diego Klabjan, and Patrick NBless · 2019
Earlier work this paper cites.
Unifying Vision-and-Language Tasks via Text Generation, May 2021
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal · 2021
Cited alongside, same era.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, Aug. 2021
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
ChartOCR: Data Extraction from Charts Images via a Deep Hybrid Framework
Junyu Luo, Zekun Li, Jinpeng Wang, and Chin-Yew Lin · 2021
Cited alongside, same era.
Towards an efficient framework for Data Extraction from Chart Images, May 2021
Weihong Ma, Hesuo Zhang, Shuang Yan, Guangshun Yao, Yichao Huang, Hui Li, Yaqiang Wu, and Lianwen Jin · 2021
Cited alongside, same era.
A comparative analysis of object detection algorithms in naturalistic driving videos
PaLI: A Jointly-Scaled Multilingual Language-Image Model, June 2023
Xi Chen, Xiao Wang, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut · 2023
Later among the works it cites.
Zhi-Qi Cheng, Qi Dai, Siyao Li, Jingdong Sun, Teruko Mitamura, and Alexander G. Hauptmann · 2023
Later among the works it cites.
LineFormer: Rethinking Line Chart Data Extraction as Instance Segmentation, May 2023
Jay Lal, Aditya Mitkari, Mahesh Bhosale, and David Doermann · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ce Zhang and Azim Eskandarian · 2021
Cited alongside, same era.
Attention-based neural network for driving environment complexity perception
Ce Zhang, Azim Eskandarian, and Xuelai Du · 2021
Cited alongside, same era.
Masked-attention Mask Transformer for Universal Image Segmentation, June 2022
Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar · 2022
Cited alongside, same era.
MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question Answering
Yang Ding, Jing Yu, Bang Liu, Yue Hu, Mingxin Cui, and Qi Wu · 2022
Cited alongside, same era.
Neighborhood Attention Transformer, Apr. 2022
Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Humphrey Shi · 2022
Cited alongside, same era.
OCR-free Document Understanding Transformer, Oct. 2022
Geewook Kim, Teakgyu Hong, Moonbin Yim, Jeongyeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park · 2022
Cited alongside, same era.
Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding, Oct. 2022
Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu, Julian Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova · 2022
Cited alongside, same era.
MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering, Dec. 2022
Fangyu Liu, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, and Julian Martin Eisenschlos · 2022
Cited alongside, same era.
Ahmed Masry, Parsa Kavehzadeh, Xuan Long Do, Enamul Hoque, and Shafiq Joty · 2023
Later among the works it cites.
LineEX: Data Extraction from Scientific Line Charts
Shivasankaran V P, Muhammad Yusuf Hassan, and Mayank Singh · 2023
Later among the works it cites.
The art of socratic questioning: Recursive thinking with large language models
Jingyuan Qi, Zhiyang Xu, Ying Shen, Minqian Liu, Di Jin, Qifan Wang, and Lifu Huang · 2023
Later among the works it cites.
DAT++: Spatially Dynamic Vision Transformer with Deformable Attention, Sept. 2023
Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao Huang · 2023
Later among the works it cites.
ChartDETR: A Multi-shape Detection Network for Visual Chart Recognition, Aug. 2023
Wenyuan Xue, Dapeng Chen, Baosheng Yu, Yifei Chen, Sai Zhou, and Wei Peng · 2023
Later among the works it cites.
Motiontrack: End-to-end transformer-based multi-object tracking with lidar-camera fusion
Ce Zhang, Chengjie Zhang, Yiluan Guo, Lingji Chen, and Michael Happold · 2023
Later among the works it cites.
Number-adaptive prototype learning for 3d point cloud semantic segmentation
Yangheng Zhao, Jun Wang, Xiaolong Li, Yue Hu, Ce Zhang, Yanfeng Wang, and Siheng Chen · 2023
Later among the works it cites.
Enhanced Chart Understanding via Visual Language Pre-training on Plot Table Pairs
Mingyang Zhou, Yi Fung, Long Chen, Christopher Thomas, Heng Ji, and Shih-Fu Chang · 2023
Later among the works it cites.
Saliendet: A saliency-based feature enhancement algorithm for object detection for autonomous driving
Ning Ding, Ce Zhang, and Azim Eskandarian · 2024
Closest in time.
Improved Baselines with Visual Instruction Tuning, May 2024
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2024
Closest in time.
Multimodal instruction tuning with conditional mixture of lora
Ying Shen, Zhiyang Xu, Qifan Wang, Yu Cheng, Wenpeng Yin, and Lifu Huang · 2024
Closest in time.
Vision-flan: Scaling human-labeled tasks in visual instruction tuning
Zhiyang Xu, Chao Feng, Rulin Shao, Trevor Ashby, Ying Shen, Di Jin, Yu Cheng, Qifan Wang, and Lifu Huang · 2024
Closest in time.