Fetching the paper…
Reading the bibliography…
One property that remains lacking in image captions generated by contemporary methods is discriminability: being able to tell two images apart given the caption for one of them.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Lex Weaver and Nigel Tao · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Wsabie: Scaling up to large vocabulary image annotation
Jason Weston, Samy Bengio, and Nicolas Usunier · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Tomas Mikolov, et al · 2013
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara L Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Baidu Research, and Alan Yuille · 2015
Earlier work this paper cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Lei Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Order-Embeddings of Images and Language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun · 2015
Earlier work this paper cites.
Visual turing test for computer vision systems
Donald Geman, Stuart Geman, Neil Hallonquist, and Laurent Younes · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Cited alongside, same era.
Self-critical sequence training for image captioning
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel · 2016
Cited alongside, same era.
Multimodal convolutional neural networks for matching image and sentence
Lin Ma, Zhengdong Lu, Lifeng Shang, and Hang Li · 2016
Cited alongside, same era.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher · 2016
Cited alongside, same era.
Review networks for caption generation
Zhilin Yang, Ye Yuan, Yuexin Wu, William W Cohen, and Ruslan R Salakhutdinov · 2016
Cited alongside, same era.
Context-aware captions from context-agnostic supervision
Ramakrishna Vedantam, Samy Bengio, Kevin Murphy, Devi Parikh, and Gal Chechik · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and vqa
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2017
Later among the works it cites.
Stack-captioning: Coarse-to-fine learning for image captioning
Jiuxiang Gu, Jianfei Cai, Gang Wang, and Tsuhan Chen · 2017
Later among the works it cites.
Skeleton key: Image captioning by skeleton-attribute decomposition
Yufei Wang, Zhe Lin, Xiaohui Shen, Scott Cohen, and Garrison W Cottrell · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention Correctness in Neural Image Captioning
Chenxi Liu, Junhua Mao, Fei Sha, and Alan Yuille · 2016
Cited alongside, same era.
Boosting image captioning with attributes
Ting Yao, Yingwei Pan, Yehao Li, Zhaofan Qiu, and Tao Mei · 2016
Cited alongside, same era.
Sequence Level Training with Recurrent Neural Networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Cited alongside, same era.
Learning Deep Structure-Preserving Image-Text Embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2016
Cited alongside, same era.
Dual attention networks for multimodal reasoning and matching
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim · 2016
Cited alongside, same era.
Reasoning About Pragmatics with Neural Listeners and Speakers
Jacob Andreas and Dan Klein · 2016
Cited alongside, same era.
Modeling Context in Referring Expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Cited alongside, same era.
Liwei Wang, Alexander Schwing, and Svetlana Lazebnik · 2017
Later among the works it cites.
Actor-critic sequence training for image captioning
Li Zhang, Flood Sung, Feng Liu, Tao Xiang, Shaogang Gong, Yongxin Yang, and Timothy M Hospedales · 2017
Later among the works it cites.
Video captioning via hierarchical reinforcement learning
Xin Wang, Wenhu Chen, Jiawei Wu, Yuan-Fang Wang, and William Yang Wang · 2017
Later among the works it cites.
Deep reinforcement learning-based image captioning with embedding reward
Zhou Ren, Xiaoyu Wang, Ning Zhang, Xutao Lv, and Li-Jia Li · 2017
Later among the works it cites.
Learning two-branch neural networks for image-text matching tasks
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2017
Later among the works it cites.
Comprehension-guided referring expressions
Ruotian Luo and Gregory Shakhnarovich · 2017
Later among the works it cites.
Learning cooperative visual dialog agents with deep reinforcement learning
Abhishek Das, Satwik Kottur, José MF Moura, Stefan Lee, and Dhruv Batra · 2017
Later among the works it cites.
Learning to disambiguate by asking discriminative questions
Yining Li, Chen Huang, Xiaoou Tang, and Chen-Change Loy · 2017
Later among the works it cites.
Contrastive learning for image captioning
Bo Dai and Dahua Lin · 2017
Later among the works it cites.
Jiasen Lu, Anitha Kannan, Jianwei Yang, Devi Parikh, and Dhruv Batra · 2017
Later among the works it cites.
Towards diverse and natural image descriptions via a conditional gan
Bo Dai, Dahua Lin, Raquel Urtasun, and Sanja Fidler · 2017
Later among the works it cites.
Speaking the same language: Matching machine to human captions by adversarial training
Rakshith Shetty, Marcus Rohrbach, Lisa Anne Hendricks, Mario Fritz, and Bernt Schiele · 2017
Later among the works it cites.
Vse++: Improved visual-semantic embeddings
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2017
Later among the works it cites.
Teaching machines to describe images via natural language feedback
Huan Ling and Sanja Fidler · 2017
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Later among the works it cites.