Fetching the paper…
Reading the bibliography…
Evaluation metrics for image captioning face two challenges.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Finding frequent items in data streams
M. Charikar, K. Chen, and M. Farach-Colton · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Earlier work this paper cites.
Fast and scalable polynomial kernels via explicit feature maps
N. Pham and R. Pagh · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. D. A. Lavie · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
http://mscoco.org/dataset/#captions-challenge2015
The coco 2015 captioning challenge · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Compact bilinear pooling
Y. Gao, O. Beijbom, N. Zhang, and T. Darrell · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Generating images with recurrent adversarial networks
D. J. Im, C. D. Kim, H. Jiang, and R. Memisevic · 2016
Later among the works it cites.
Show, adapt and tell: Adversarial training of cross-domain image captioner
T.-H. Chen, Y.-H. Liao, C.-Y. Chuang, W.-T. Hsu, J. Fu, and M. Sun · 2017
Later among the works it cites.
Towards diverse and natural image descriptions via a conditional gan
B. Dai, S. Fidler, R. Urtasun, and D. Lin · 2017
Later among the works it cites.
Adversarial evaluation of dialogue models
A. Kannan and O. Vinyals · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al · 2016
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Cited alongside, same era.
Generating sentences from a continuous space
S. R. Bowman, L. Vilnis, O. Vinyals, A. M. Dai, R. Jozefowicz, and S. Bengio · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Adversarial learning for neural dialogue generation
J. Li, W. Monroe, T. Shi, A. Ritter, and D. Jurafsky · 2017
Later among the works it cites.
Recurrent topic-transition gan for visual paragraph generation
X. Liang, Z. Hu, H. Zhang, C. Gan, and E. P. Xing · 2017
Later among the works it cites.
Improved image captioning via policy gradient optimization of spider
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy · 2017
Later among the works it cites.
Towards an automatic turing test: Learning to evaluate dialogue responses
R. Lowe, M. Noseworthy, I. Serban, N. Angelard-Gontier, Y. Bengio, and J. Pineau · 2017
Later among the works it cites.
Speaking the same language: Matching machine to human captions by adversarial training
R. Shetty, M. Rohrbach, L. A. Hendricks, M. Fritz, and B. Schiele · 2017
Later among the works it cites.