Fetching the paper…
Reading the bibliography…
Existing synthetic datasets (FigureQA, DVQA) for reasoning over plots do not contain variability in data labels, real-valued data, or complex reasoning questions.
An overview of the tesseract ocr engine
R. Smith · 2007
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
VQA: visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
R. B. Girshick · 2015
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
P. Pasupat and P. Liang · 2015
Earlier work this paper cites.
Image question answering: A visual semantic embedding model and a new dataset
M. Ren, R. Kiros, and R. S. Zemel · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
S. Ren, K. He, R. B. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. J. Seo, H. Hajishirzi, and A. Farhadi · 2016
Earlier work this paper cites.
SSD: single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. E. Reed, C. Fu, and A. C. Berg · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Learning a natural language interface with neural programmer
A. Neelakantan, Q. V. Le, M. Abadi, A. McCallum, and D. Amodei · 2016
Cited alongside, same era.
Training recurrent answering units with joint loss minimization for VQA
H. Noh and B. Han · 2016
Cited alongside, same era.
Figureseer: Parsing result-figures in research papers
N. Siegel, Z. Horvitz, R. Levin, S. K. Divvala, and A. Farhadi · 2016
Cited alongside, same era.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
A. Kembhavi, M. J. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi · 2017
Later among the works it cites.
Hadamard product for low-rank bilinear pooling
J. Kim, K. W. On, W. Lim, J. Kim, J. Ha, and B. Zhang · 2017
Later among the works it cites.
Neural semantic parsing with type constraints for semi-structured tables
J. Krishnamurthy, P. Dasigi, and M. Gardner · 2017
Later among the works it cites.
Feature pyramid networks for object detection
T. Lin, P. Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie · 2017
Later among the works it cites.
YOLO9000: better, faster, stronger
J. Redmon and A. Farhadi · 2017
Later among the works it cites.
A corpus of natural language for visual reasoning
A. Suhr, M. Lewis, J. Yeh, and Y. Artzi · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Yang, X. He, J. Gao, L. Deng, and A. J. Smola · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. J. Smola · 2016
Cited alongside, same era.
Scatteract: Automated extraction of data from scatter plots
M. Cliche, D. S. Rosenberg, D. Madeka, and C. Yee · 2017
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Cited alongside, same era.
Mask R-CNN
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick · 2017
Cited alongside, same era.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. B. Girshick · 2017
Cited alongside, same era.
Figureqa: An annotated figure dataset for visual reasoning
S. E. Kahou, A. Atkinson, V. Michalski, Á. Kádár, A. Trischler, and Y. Bengio · 2017
Cited alongside, same era.
Later among the works it cites.
Detectron
R. Girshick, I. Radosavovic, G. Gkioxari, P. Dollár, and K. He · 2018
Later among the works it cites.
Neural multi-step reasoning for question answering on semi-structured tables
T. Haug, O. Ganea, and P. Grnarova · 2018
Later among the works it cites.
DVQA: understanding data visualizations via question answering
K. Kafle, S. Cohen, B. L. Price, and C. Kanan · 2018
Later among the works it cites.
Bilinear attention networks
J. Kim, J. Jun, and B. Zhang · 2018
Later among the works it cites.
Pythia-a platform for vision & language research
A. Singh, V. Natarajan, Y. Jiang, X. Chen, M. Shah, M. Rohrbach, D. Batra, and D. Parikh · 2018
Later among the works it cites.
Towards vqa models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Closest in time.