Fetching the paper…
Reading the bibliography…
Despite the substantial progress in recent years, the image captioning techniques are still far from being perfect.Sentences produced by existing methods, e.g.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi · 2011
Earlier work this paper cites.
Collective generation of natural image descriptions
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2012
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. D. A. Lavie · 2014
Earlier work this paper cites.
Simple image description generator via a linear phrase-based approach
R. Lebret, P. O. Pinheiro, and R. Collobert · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
M. Mirza and S. Osindero · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Cited alongside, same era.
Language models for image captioning: The quirks and what works
J. Devlin, H. Cheng, H. Fang, S. Gupta, L. Deng, X. He, G. Zweig, and M. Mitchell · 2015
Cited alongside, same era.
Optimization of image description metrics using policy gradient methods
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy · 2016
Later among the works it cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
J. Lu, C. Xiong, D. Parikh, and R. Socher · 2016
Later among the works it cites.
Generative adversarial text to image synthesis
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee · 2016
Later among the works it cites.
Self-critical sequence training for image captioning
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel · 2016
Later among the works it cites.
Review networks for caption generation
Z. Yang, Y. Yuan, Y. Wu, W. W. Cohen, and R. R. Salakhutdinov · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Devlin, S. Gupta, R. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, et al · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros · 2016
Cited alongside, same era.
Image captioning with semantic attention
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Later among the works it cites.
Seqgan: sequence generative adversarial nets with policy gradient
L. Yu, W. Zhang, J. Wang, and Y. Yu · 2016
Later among the works it cites.
Image caption generation with text-conditional semantic attention
L. Zhou, C. Xu, P. Koch, and J. J. Corso · 2016
Later among the works it cites.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Closest in time.
A hierarchical approach for generating descriptive image paragraphs
J. Krause, J. Johnson, R. Krishna, and L. Fei-Fei · 2017
Closest in time.
Scene graph generation from objects, phrases and region captions
Y. Li, W. Ouyang, B. Zhou, K. Wang, and X. Wang · 2017
Closest in time.