Fetching the paper…
Reading the bibliography…
This paper presents a new metric called TIGEr for the automatic evaluation of image captioning systems.
No metrics are perfect: Adversarial reward learning for visual storytelling
Xin Wang, Wenhu Chen, Yuan-Fang Wang, and William Yang Wang. 2018b · 1909
Earlier work this paper cites.
On information and sufficiency
Solomon Kullback and Richard A Leibler. 1951 · 1951
Earlier work this paper cites.
A machine learning approach to the automatic evaluation of machine translation
Simon Corston-Oliver, Michael Gamon, and Chris Brockett. 2001 · 2001
Earlier work this paper cites.
Cumulated gain-based evaluation of ir techniques
Kalervo Järvelin and Jaana Kekäläinen. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
A learning approach to improving sentence-level mt evaluation
Alex Kulesza and Stuart M Shieber. 2004 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Tomas Mikolov, et al. 2013 · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013 · 2013
Earlier work this paper cites.
Comparing automatic evaluation measures for image description
Desmond Elliott and Frank Keller. 2014 · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li F Fei-Fei. 2014 · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2014 · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al. 2015 · 2015
Cited alongside, same era.
Revisiting visual question answering baselines
Allan Jabri, Armand Joulin, and Laurens van der Maaten. 2016 · 2016
Later among the works it cites.
Towards diverse and natural image descriptions via a conditional gan
Bo Dai, Sanja Fidler, Raquel Urtasun, and Dahua Lin. 2017 · 2017
Later among the works it cites.
Semantic compositional networks for visual captioning
Zhe Gan, Chuang Gan, Xiaodong He, Yunchen Pu, Kenneth Tran, Jianfeng Gao, Lawrence Carin, and Li Deng. 2017 · 2017
Later among the works it cites.
Towards an automatic Turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and vqa
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li. 2015 · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Cited alongside, same era.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
Raffaella Bernardi, Ruket Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, and Barbara Plank. 2016 · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016 · 2016
Cited alongside, same era.
From images to sentences through scene description graphs using commonsense reasoning and knowledge
Somak Aditya, Yezhou Yang, Chitta Baral, Cornelia Fermuller, and Yiannis Aloimonos. 2015a
Cited in the paper.
Hongge Chen, Huan Zhang, Pin-Yu Chen, Jinfeng Yi, and Cho-Jui Hsieh. 2018 · 2018
Later among the works it cites.
Learning to evaluate image captioning
Yin Cui, Guandao Yang, Andreas Veit, Xun Huang, and Serge Belongie. 2018 · 2018
Later among the works it cites.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018 · 2018
Later among the works it cites.
Neural approaches to conversational ai
Jianfeng Gao, Michel Galley, and Lihong Li. 2019 · 2019
Closest in time.
Hierarchically structured reinforcement learning for topically coherent visual story generation
Qiuyuan Huang, Zhe Gan, Asli Celikyilmaz, Dapeng Wu, Jianfeng Wang, and Xiaodong He. 2019 · 2019
Closest in time.