Fetching the paper…
Reading the bibliography…
The aim of image captioning is to generate textual description of a given image.
Tree-structured neural machine for linguistics-aware sentence generation
Zhou, G.; Luo, P.; Cao, R.; Xiao, Y.; Lin, F.; Chen, B.; and He, Q. 2018 · 2011
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Hodosh, M.; Young, P.; and Hockenmaier, J. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Show and Tell: A Neural Image Caption Generator
Vinyals, O.; Toshev, A.; Bengio, S.; and Erhan, D. 2014 · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P.; Lai, A.; Hodosh, M.; and Hockenmaier, J. 2014 · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A.; and Fei-Fei, L. 2015 · 2015
Cited alongside, same era.
Image-to-Tree: A Tree-Structured Decoder for Image Captioning
Ma, Z.; Yuan, C.; Cheng, Y.; and Zhu, X. 2019 · 2019
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A.; Salakhudinov, R.; Zemel, R.; and Bengio, Y. 2015 · 2057
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…