Fetching the paper…
Reading the bibliography…
Deep models are the defacto standard in visual decision models due to their impressive performance on a wide array of visual tasks.
A model of inexact reasoning in medicine
E. H. Shortliffe and B. G. Buchanan · 1975
Earlier work this paper cites.
Reconstructive expert system explanation
M. R. Wick and W. B. Thompson · 1992
Earlier work this paper cites.
A metric for distributions with applications to image databases
Y. Rubner, C. Tomasi, and L. J. Guibas · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: a package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
An explainable artificial intelligence system for small-unit tactical behavior
M. Van Lent, W. Fisher, and M. Mancuso · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Explainable artificial intelligence for training and tutoring
H. C. Lane, M. G. Core, M. Van Lent, S. Solomon, and D. Gomboc · 2005
Earlier work this paper cites.
Building explainable artificial intelligence systems
M. G. Core, H. C. Lane, M. Van Lent, D. Gomboc, S. Solomon, and M. Rosenberg · 2006
Earlier work this paper cites.
Fast and robust earth mover’s distances
O. Pele and M. Werman · 2009
Earlier work this paper cites.
Caltech-UCSD Birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona · 2010
Earlier work this paper cites.
What makes paris look like paris?
C. Doersch, S. Singh, A. Gupta, J. Sivic, and A. Efros · 2012
Earlier work this paper cites.
Opensurfaces: A richly annotated catalog of surface appearance
S. Bell, P. Upchurch, N. Snavely, and K. Bala · 2013
Earlier work this paper cites.
How do you tell a blackbird from a crow?
T. Berg and P. N. Belhumeur · 2013
Earlier work this paper cites.
2d human pose estimation: New benchmark and state of the art analysis
M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele · 2014
Cited alongside, same era.
Justification narratives for individual classifications
O. Biran and K. McKeown · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Fine-grained activity recognition with holistic and pose based features
L. Pishchulin, M. Andriluka, and B. Schiele · 2014
Cited alongside, same era.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Cited alongside, same era.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
A. Das, H. Agrawal, C. L. Zitnick, D. Parikh, and D. Batra · 2016
Closest in time.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, M. Rohrbach, S. Venugopalan, S. Guadarrama, K. Saenko, and T. Darrell · 2016
Closest in time.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Closest in time.
Compact bilinear pooling
Y. Gao, O. Beijbom, N. Zhang, and T. Darrell · 2016
Closest in time.
Generating visual explanations
L. A. Hendricks, Z. Akata, M. Rohrbach, J. Donahue, B. Schiele, and T. Darrell · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
Abc-cnn: An attention based convolutional neural network for visual question answering
K. Chen, J. Wang, L.-C. Chen, H. Gao, W. Xu, and R. Nevatia · 2015
Cited alongside, same era.
On the relationship between visual attributes and convolutional networks
V. Escorcia, J. C. Niebles, and B. Ghanem · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
J. Kim, K. W. On, J. Kim, J. Ha, and B. Zhang · 2016
Closest in time.
Learning models for actions and person-object interactions with transfer to question answering
A. Mallya and S. Lazebnik · 2016
Closest in time.
Learning deep representations of fine-grained visual descriptions
S. Reed, Z. Akata, H. Lee, and B. Schiele · 2016
Closest in time.
Learning what and where to draw
S. Reed, Z. Akata, S. Mohan, S. Tenka, B. Schiele, and H. Lee · 2016
Closest in time.
Where to look: Focus regions for visual question answering
K. J. Shih, S. Singh, and D. Hoiem · 2016
Closest in time.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Closest in time.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Closest in time.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Closest in time.
Visual7W: Grounded Question Answering in Images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Closest in time.