Fetching the paper…
Reading the bibliography…
Bar charts are an effective way to convey numeric information, but today's algorithms cannot parse them.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
The automated understanding of simple bar charts
S. Elzer, S. Carberry, and I. Zukerman · 2011
Earlier work this paper cites.
Revision: Automated classification, analysis and redesign of chart images
M. Savva, N. Kong, A. Chhajta, L. Fei-Fei, M. Agrawala, and J. Heer · 2011
Earlier work this paper cites.
Extraction and interpretation of charts in technical documents
J. S. Kallimani, K. Srinivasa, and R. B. Eswara · 2013
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Automatic extraction of data from bar charts
R. A. Al-Zaidy and C. L. Giles · 2015
Earlier work this paper cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Earlier work this paper cites.
Deep compositional question answering with neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Cited alongside, same era.
Answer-type prediction for visual question answering
K. Kafle and C. Kanan · 2016
Cited alongside, same era.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi · 2016
Cited alongside, same era.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2016
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Later among the works it cites.
An analysis of visual question answering algorithms
K. Kafle and C. Kanan · 2017
Later among the works it cites.
Visual question answering: Datasets, algorithms, and future challenges
K. Kafle and C. Kanan · 2017
Later among the works it cites.
Data augmentation for visual question answering
K. Kafle, M. Yousefhussien, and C. Kanan · 2017
Later among the works it cites.
Figureqa: An annotated figure dataset for visual reasoning
S. E. Kahou, A. Atkinson, V. Michalski, A. Kadar, A. Trischler, and Y. Bengio · 2017
Later among the works it cites.
Show, ask, attend, and answer: A strong baseline for visual question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Training recurrent answering units with joint loss minimization for VQA
H. Noh and B. Han · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. J. Smola · 2016
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Cited alongside, same era.
Learning to reason: End-to-end module networks for visual question answering
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko · 2017
Cited alongside, same era.
V. Kazemi and A. Elqursh · 2017
Later among the works it cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
A. Kembhavi, M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi · 2017
Later among the works it cites.
Hadamard product for low-rank bilinear pooling
J.-H. Kim, K.-W. On, J. Kim, J.-W. Ha, and B.-T. Zhang · 2017
Later among the works it cites.
Reverse-engineering visualizations: Recovering visual encodings from chart images
J. Poco and J. Heer · 2017
Later among the works it cites.
Visual question answering: A survey of methods and datasets
Q. Wu, D. Teney, P. Wang, C. Shen, A. Dick, and A. v. d. Hengel · 2017
Later among the works it cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi · 2018
Closest in time.