Fetching the paper…
Reading the bibliography…
Image caption generation is a long standing and challenging problem at the intersection of computer vision and natural language processing.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
C. H. Lampert, H. Nickisch, and S. Harmeling · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara L Berg · 2011
Earlier work this paper cites.
Evaluating knowledge transfer and zero-shot learning in a large-scale setting
Marcus Rohrbach, Michael Stark, and Bernt Schiele · 2011
Earlier work this paper cites.
Label-embedding for attribute-based classification
Zeynep Akata, Florent Perronnin, Zaid Harchaoui, and Cordelia Schmid · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier · 2013
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
Girish Kulkarni, Visruth Premraj, Vicente Ordonez, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg · 2013
Earlier work this paper cites.
Large-scale object classification using label relation graphs
Jia Deng, Nan Ding, Yangqing Jia, Andrea Frome, Kevin Murphy, Samy Bengio, Yuan Li, Hartmut Neven, and Hartwig Adam · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Lsda: Large scale detection through adaptation
Judy Hoffman, Sergio Guadarrama, Eric S Tzeng, Ronghang Hu, Jeff Donahue, Ross Girshick, Trevor Darrell, and Kate Saenko · 2014
Earlier work this paper cites.
Zero-shot recognition with unreliable attributes
Dinesh Jayaraman and Kristen Grauman · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Overfeat: Integrated recognition, localization and detection using convolutional networks
Pierre Sermanet, David Eigen, Xiang Zhang, Michaël Mathieu, Robert Fergus, and Yann Lecun · 2014
Earlier work this paper cites.
The fastest deformable part model for object detection
Junjie Yan, Zhen Lei, Longyin Wen, and Stan Li · 2014
Earlier work this paper cites.
Evaluation of output embeddings for fine-grained image classification
Zeynep Akata, Scott Reed, Daniel Walter, Honglak Lee, and Bernt Schiele · 2015
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Spatial pyramid pooling in deep convolutional networks for visual recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Detector discovery in the wild: Joint multiple instance and representation learning
Judy Hoffman, Deepak Pathak, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Cited alongside, same era.
Automatic concept discovery from parallel text and visual corpora
Chen Sun, Chuang Gan, and Ram Nevatia · 2015
Cited alongside, same era.
Pointer networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Later among the works it cites.
Yolo9000: Better, faster, stronger
Joseph Redmon and Ali Farhadi · 2017
Later among the works it cites.
Dsod: Learning deeply supervised object detectors from scratch
Zhiqiang Shen, Zhuang Liu, Jianguo Li, Yu-Gang Jiang, Yurong Chen, and Xiangyang Xue · 2017
Later among the works it cites.
Captioning images with diverse objects
Subhashini Venugopalan, Lisa Anne Hendricks, Marcus Rohrbach, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2017
Later among the works it cites.
Dense captioning with joint inference and visual context
Linjie Yang, Kevin Tang, Jianchao Yang, and Li-Jia Li · 2017
Later among the works it cites.
Incorporating copying mechanism in image captioning for learning novel objects
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Cited alongside, same era.
Deep compositional captioning: Describing novel object categories without paired training data
Lisa Anne Hendricks, Subhashini Venugopalan, Marcus Rohrbach, Raymond Mooney, Kate Saenko, and Trevor Darrell · 2016
Cited alongside, same era.
Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks
Sean Bell, C Lawrence Zitnick, Kavita Bala, and Ross Girshick · 2016
Cited alongside, same era.
Weakly supervised deep detection networks
Hakan Bilen and Andrea Vedaldi · 2016
Cited alongside, same era.
Synthesized classifiers for zero-shot learning
Soravit Changpinyo, Wei-Lun Chao, Boqing Gong, and Fei Sha · 2016
Cited alongside, same era.
Weakly supervised object localization with multi-fold multiple instance learning
Ramazan Gokberk Cinbis, Jakob Verbeek, and Cordelia Schmid · 2016
Cited alongside, same era.
Later among the works it cites.
Obj2text: Generating visually descriptive language from object layouts
Xuwang Yin and Vicente Ordonez · 2017
Later among the works it cites.
Zero-shot object detection
Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran · 2018
Later among the works it cites.
Zero-shot object detection by hybrid region embedding
Berkan Demirel, Ramazan Gokberk Cinbis, and Nazli Ikizler-Cinbis · 2018
Later among the works it cites.
Cornernet: Detecting objects as paired keypoints
Hei Law and Jia Deng · 2018
Later among the works it cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priyal Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2018
Later among the works it cites.
Zero-shot learning using synthesised unseen visual data with diffusion regularisation
Yang Long, Li Liu, Fumin Shen, Ling Shao, and Xuelong Li · 2018
Later among the works it cites.
Neural baby talk
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2018
Later among the works it cites.
Zero-shot learning via attribute regression and class prototype rectification
Changzhi Luo, Zhetao Li, Kaizhu Huang, Jiashi Feng, and Meng Wang · 2018
Later among the works it cites.
A generative model for zero shot learning using conditional variational autoencoders
Ashish Mishra, Shiva Krishna Reddy, Anurag Mittal, and Hema A Murthy · 2018
Later among the works it cites.
Transductive unbiased embedding for zero-shot learning
Jie Song, Chengchao Shen, Yezhou Yang, Yang Liu, and Mingli Song · 2018
Later among the works it cites.
Zero-shot learning via class-conditioned deep generative models
Wenlin Wang, Yunchen Pu, Vinay Kumar Verma, Kai Fan, Yizhe Zhang, Changyou Chen, Piyush Rai, and Lawrence Carin · 2018
Later among the works it cites.
Decoupled novel object captioner
Yu Wu, Linchao Zhu, Lu Jiang, and Yi Yang · 2018
Later among the works it cites.
Zero-shot learning-a comprehensive evaluation of the good, the bad and the ugly
Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata · 2018
Later among the works it cites.
Zero-shot learning via latent space encoding
Yunlong Yu, Zhong Ji, Jichang Guo, and Zhongfei Zhang · 2018
Later among the works it cites.
A generative adversarial approach for zero-shot learning from noisy texts
Yizhe Zhu, Mohamed Elhoseiny, Bingchen Liu, Xi Peng, and Ahmed Elgammal · 2018
Later among the works it cites.
Gradient matching generative networks for zero-shot learning
Mert Bulent Sariyildiz and Ramazan Gokberk Cinbis · 2019
Closest in time.