Fetching the paper…
Reading the bibliography…
Many vision-language tasks can be reduced to the problem of sequence prediction for natural language output.
Toward memory-based reasoning
Craig Stanfill and David Waltz. 1986 · 1986
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 1998 · 1998
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation. In ACL
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network. In Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology-Volume 1
Kristina Toutanova, Dan Klein, Christopher D Manning, and Yoram Singer. 2003 · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks. In NIPS
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context. In ECCV
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering. In ICCV
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate. In ICLR
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John Platt, et al · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions. In CVPR
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2015 · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks. In NIPS
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and VQA
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2017 · 2017
Later among the works it cites.
Structcap: Structured semantic embedding for image captioning. In Proceedings of the 2017 ACM on Multimedia Conference . ACM, 46–54
Fuhai Chen, Rongrong Ji, Jinsong Su, Yongjian Wu, and Yunsheng Wu. 2017a · 2017
Later among the works it cites.
Visual dialog. In CVPR
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Later among the works it cites.
Stack-captioning: Coarse-to-fine learning for image captioning
Jiuxiang Gu, Jianfei Cai, Gang Wang, and Tsuhan Chen. 2017 · 2017
Later among the works it cites.
Learning to Reason: End-To-End Module Networks for Visual Question Answering. In ICCV
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cider: Consensus-based image description evaluation. In CVPR
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator. In CVPR
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention. In ICML
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation. In ECCV
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition. In CVPR
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Revisiting visual question answering baselines. In ECCV
Allan Jabri, Armand Joulin, and Laurens van der Maaten. 2016 · 2016
Cited alongside, same era.
Improved image captioning via policy gradient optimization of spider
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Inferring and Executing Programs for Visual Reasoning. In ICCV
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Judy Hoffman, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Later among the works it cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. 2017 · 2017
Later among the works it cites.
Deep reinforcement learning-based image captioning with embedding reward
Zhou Ren, Xiaoyu Wang, Ning Zhang, Xutao Lv, and Li-Jia Li. 2017 · 2017
Later among the works it cites.
Scene graph generation by iterative message passing. In CVPR
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei. 2017 · 2017
Later among the works it cites.
SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient.. In AAAI
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu. 2017 · 2017
Later among the works it cites.
Actor-Critic Sequence Training for Image Captioning
Li Zhang, Flood Sung, Feng Liu, Tao Xiang, Shaogang Gong, Yongxin Yang, and Timothy M Hospedales. 2017b · 2017
Later among the works it cites.
Grounding Referring Expressions in Images by Variational Context. In CVPR
Hanwang Zhang, Yulei Niu, and Shih-Fu Chang. 2018 · 2018
Closest in time.