Fetching the paper…
Reading the bibliography…
Image captioning is a challenging task where the machine automatically describes an image by sentences or phrases.
Introduction to WordNet: An on-line lexical database
George A Miller, Richard Beckwith, Christiane Fellbaum, Derek Gross, and Katherine J Miller. 1990 · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In ACL-W . 65–72
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images. In ECCV . 15–29
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth. 2010 · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs. In NIPS . 1143–1151
Vicente Ordonez, Girish Kulkarni, and Tamara L Berg. 2011 · 2011
Earlier work this paper cites.
Evaluating knowledge transfer and zero-shot learning in a large-scale setting. In CVPR . 1641–1648
Marcus Rohrbach, Michael Stark, and Bernt Schiele. 2011 · 2011
Earlier work this paper cites.
Midge: Generating Image Descriptions From Computer Vision Detections. In EACL . 747–756
Margaret Mitchell, Xufeng Han, Jesse Dodge, Alyssa Mensch, Amit Goyal, Alex Berg, Kota Yamaguchi, Tamara Berg, Karl Stratos, and Hal Daumé III. 2012 · 2012
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
Girish Kulkarni, Visruth Premraj, Vicente Ordonez, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg. 2013 · 2013
Earlier work this paper cites.
Multimodal neural language models. In ICML . 595–603
Ryan Kiros, Ruslan Salakhutdinov, and Rich Zemel. 2014 · 2014
Earlier work this paper cites.
Attribute-based classification for zero-shot visual object categorization
Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context. In ECCV . 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks. In NIPS . 1171–1179
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description. In CVPR . 2625–2634
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions. In CVPR . 3128–3137
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization. In ICLR
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille. 2015b · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks. In NIPS . 91–99
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition. In ICLR
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
Sequence to sequence-video to text. In ICCV . 4534–4542
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko. 2015 · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator. In CVPR . 3156–3164
Guided open vocabulary image captioning with constrained beam search. In EMNLP
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2017 · 2017
Later among the works it cites.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. In ICML . 1126–1135
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Later among the works it cites.
Speed/accuracy trade-offs for modern convolutional object detectors. In CVPR
Jonathan Huang, Vivek Rathod, Chen Sun, Menglong Zhu, Anoop Korattikara, Alireza Fathi, Ian Fischer, Zbigniew Wojna, Yang Song, Sergio Guadarrama, et al · 2017
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning. In AAAI
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. 2017 · 2017
Later among the works it cites.
Paying Attention to Descriptions Generated by Image Captioning Models. In ICCV . 2506–2515
Hamed R Tavakoliy, Rakshith Shetty, Ali Borji, and Jorma Laaksonen. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
TensorFlow: A System for Large-Scale Machine Learning.. In OSDI , Vol. 16. 265–283
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Deep compositional captioning: Describing novel object categories without paired training data. In CVPR
Lisa Anne Henzdricks, Subhashini Venugopalan, Marcus Rohrbach, Raymond Mooney, Kate Saenko, Trevor Darrell, Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, et al · 2016
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning. In CVPR . 4565–4574
Justin Johnson, Andrej Karpathy, and Li Fei-Fei. 2016 · 2016
Cited alongside, same era.
Sequence level training with recurrent neural networks. In ICLR
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
One-shot learning with memory-augmented neural networks
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap. 2016 · 2016
Cited alongside, same era.
Matching networks for one shot learning. In NIPS . 3630–3638
Oriol Vinyals, Charles Blundell, Tim Lillicrap, Daan Wierstra, et al · 2016
Cited alongside, same era.
Captioning Images with Diverse Objects. In CVPR
Subhashini Venugopalan, Lisa Anne Hendricks, Marcus Rohrbach, Raymond Mooney, Trevor Darrell, and Kate Saenko. 2017 · 2017
Later among the works it cites.
Show and Tell: Lessons Learned from the 2015 MSCOCO Image Captioning Challenge
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. 2017 · 2017
Later among the works it cites.
Incorporating copying mechanism in image captioning for learning novel objects. In CVPR . 5263–5271
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. 2017 · 2017
Later among the works it cites.
Uncovering the Temporal Context for Video Question Answering
Linchao Zhu, Zhongwen Xu, Yi Yang, and Alexander G. Hauptmann. 2017 · 2017
Later among the works it cites.
Fast Parameter Adaptation for Few-shot Image Captioning and Visual Question Answering. In ACM on Multimedia
Xuanyi Dong, Linchao Zhu, De Zhang, Yi Yang, and Fei Wu. 2018 · 2018
Closest in time.
Neural Baby Talk. In CVPR . 7219–7228
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2018 · 2018
Closest in time.
Zero-Shot Learning - A Comprehensive Evaluation of the Good, the Bad and the Ugly
Y. Xian, C. H. Lampert, B. Schiele, and Z. Akata. 2018 · 2018
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention. In ICML . 2048–2057
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.