Fetching the paper…
Reading the bibliography…
Image captioning is an ambiguous problem, with many suitable captions for an image.
Bleu: A method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: a package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Solving the problem of cascading errors: Approximate bayesian inference for linguistic annotation pipelines
J. R. Finkel, C. D. Manning, and A. Y. Ng · 2006
Earlier work this paper cites.
A systematic exploration of diversity in machine translation
K. Gimpel, D. Batra, C. Dyer, and G. Shakhnarovich · 2013
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Exploring nearest neighbor approaches for image captioning
J. Devlin, S. Gupta, R. B. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. L. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Cited alongside, same era.
Improved image captioning via policy gradient optimization of spider
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy · 2016
Cited alongside, same era.
Situation recognition: Visual semantic role labeling for image understanding
Show and tell: Lessons learned from the 2015 mscoco image captioning challenge
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2017
Later among the works it cites.
Learning two-branch neural networks for image-text matching tasks
L. Wang, Y. Li, and S. Lazebnik · 2017
Later among the works it cites.
Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space
L. Wang, A. G. Schwing, and S. Lazebnik · 2017
Later among the works it cites.
Boosting image captioning with attributes
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei · 2017
Later among the works it cites.
Convolutional image captioning
J. Aneja, A. Deshpande, and A. Schwing · 2018
Closest in time.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Yatskar, L. Zettlemoyer, and A. Farhadi · 2016
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2017
Cited alongside, same era.
Towards diverse and natural image descriptions via a conditional gan
B. Dai, S. Fidler, R. Urtasun, and D. Lin · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin · 2017
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
E. Jang, S. Gu, and B. Poole · 2017
Cited alongside, same era.
Self-critical sequence training for image captioning
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel · 2017
Cited alongside, same era.
Speaking the same language: Matching machine to human captions by adversarial training
R. Shetty, M. Rohrbach, L. A. Hendricks, M. Fritz, and B. Schiele · 2017
Cited alongside, same era.
S. Bai, J. Z. Kolter, and V. Koltun · 2018
Closest in time.
Generating diverse and accurate visual captions by comparative adversarial learning
D. Li, X. He, Q. Huang, M.-T. Sun, and L. Zhang · 2018
Closest in time.
Neural baby talk, 2018
J. Lu, J. Yang, D. Batra, and D. Parikh · 2018
Closest in time.
Discriminability objective for training descriptive captions
R. Luo, B. Price, S. Cohen, and G. Shakhnarovich · 2018
Closest in time.
Diverse beam search for improved description of complex scenes
A. K. Vijayakumar, M. Cogswell, R. R. Selvaraju, Q. Sun, S. Lee, D. J. Crandall, and D. Batra · 2018
Closest in time.
Object counts! bringing explicit detections back into image captioning
J. Wang, P. S. Madhyastha, and L. Specia · 2018
Closest in time.
Cnn+cnn: Convolutional decoders for image captioning
Q. Wang and A. B. Chan · 2018
Closest in time.