Fetching the paper…
Reading the bibliography…
Existing models which generate textual explanations enforce task relevance through a discriminative term loss function, but such mechanisms only weakly constrain mentioned object parts to actually be present in the image.
An analysis of physician attitudes regarding computer-based clinical consultation systems
R. L. Teach and E. H. Shortliffe · 1981
Earlier work this paper cites.
Caltech-ucsd birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona · 2010
Earlier work this paper cites.
Relative attributes
D. Parikh and K. Grauman · 2011
Earlier work this paper cites.
Justification narratives for individual classifications
O. Biran and K. McKeown · 2014
Cited alongside, same era.
Generating visual explanations
L. A. Hendricks, Z. Akata, M. Rohrbach, J. Donahue, B. Schiele, and T. Darrell · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Cited alongside, same era.
Attentive explanations: Justifying decisions and pointing to the evidence
D. H. Park, L. A. Hendricks, Z. Akata, B. Schiele, T. Darrell, and M. Rohrbach · 2016
Later among the works it cites.
Learning deep representations of fine-grained visual descriptions
S. Reed, Z. Akata, H. Lee, and B. Schiele · 2016
Later among the works it cites.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…