Fetching the paper…
Reading the bibliography…
It is highly desirable yet challenging to generate image captions that can describe novel objects which are unseen in caption-labeled training data, a capability that is evaluated in the novel object captioning challenge (nocaps).
Visualizing data using t-SNE
Maaten, L. v. d.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Farhadi, A.; Hejrati, M.; Sadeghi, M. A.; Young, P.; Rashtchian, C.; Hockenmaier, J.; and Forsyth, D. 2010 · 2010
Earlier work this paper cites.
Corpus-guided sentence generation of natural images
Yang, Y.; Teo, C.; Daumé III, H.; and Aloimonos, Y. 2011 · 2011
Earlier work this paper cites.
Collective generation of natural image descriptions
Kuznetsova, P.; Ordonez, V.; Berg, A.; Berg, T.; and Choi, Y. 2012 · 2012
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
Mitchell, M.; Dodge, J.; Goyal, A.; Yamaguchi, K.; Stratos, K.; Han, X.; Mensch, A.; Berg, A.; Berg, T.; and Daumé III, H. 2012 · 2012
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
Kulkarni, G.; Premraj, V.; Ordonez, V.; Dhar, S.; Li, S.; Choi, Y.; Berg, A. C.; and Berg, T. L. 2013 · 2013
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P.; Lai, A.; Hodosh, M.; and Hockenmaier, J. 2014 · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X.; Fang, H.; Lin, T.-Y.; Vedantam, R.; Gupta, S.; Dollár, P.; and Zitnick, C. L. 2015 · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Fang, H.; Gupta, S.; Iandola, F.; Srivastava, R. K.; Deng, L.; Dollár, P.; Gao, J.; He, X.; Mitchell, M.; Platt, J. C.; et al. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Deep compositional captioning: Describing novel object categories without paired training data
Hendricks, L. A.; Venugopalan, S.; Rohrbach, M.; Mooney, R.; Saenko, K.; and Darrell, T. 2016 · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Johnson, J.; Karpathy, A.; and Fei-Fei, L. 2016 · 2016
Earlier work this paper cites.
End-to-end people detection in crowded scenes
Stewart, R.; Andriluka, M.; and Ng, A. Y. 2016 · 2016
Earlier work this paper cites.
Rich image captioning in the wild
Tran, K.; He, X.; Zhang, L.; Sun, J.; Carapcea, C.; Thrasher, C.; Buehler, C.; and Sienkiewicz, C. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y.; Schuster, M.; Chen, Z.; Le, Q. V.; Norouzi, M.; Macherey, W.; Krikun, M.; Cao, Y.; Gao, Q.; Macherey, K.; et al. 2016 · 2016
Earlier work this paper cites.
Guided open vocabulary image captioning with constrained beam search
Anderson, P.; Fernando, B.; Johnson, M.; and Gould, S. 2017 · 2017
Earlier work this paper cites.
Mask R-CNN
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. 2017 · 2017
Cited alongside, same era.
Self-critical sequence training for image captioning
Rennie, S. J.; Marcheret, E.; Mroueh, Y.; Ross, J.; and Goel, V. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Captioning images with diverse objects
Venugopalan, S.; Anne Hendricks, L.; Rohrbach, M.; Mooney, R.; Darrell, T.; and Saenko, K. 2017 · 2017
Cited alongside, same era.
Incorporating copying mechanism in image captioning for learning novel objects
Yao, T.; Pan, Y.; Li, Y.; and Mei, T. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Later among the works it cites.
Objects365: A large-scale, high-quality dataset for object detection
Shao, S.; Li, Z.; Zhang, T.; Peng, C.; Yu, G.; Zhang, X.; Li, J.; and Sun, J. 2019 · 2019
Later among the works it cites.
Connecting Language to Images: A Progressive Attention-Guided Network for Simultaneous Image Captioning and Language Grounding
Song, L.; Liu, J.; Qian, B.; and Chen, Y. 2019 · 2019
Later among the works it cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Su, W.; Zhu, X.; Cao, Y.; Li, B.; Lu, L.; Wei, F.; and Dai, J. 2019 · 2019
Later among the works it cites.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Tan, H.; and Bansal, M. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Neural baby talk
Lu, J.; Yang, J.; Batra, D.; and Parikh, D. 2018 · 2018
Cited alongside, same era.
Improving Language Understanding by Generative Pre-Training
Radford, A. 2018 · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Cited alongside, same era.
Decoupled novel object captioner
Wu, Y.; Zhu, L.; Jiang, L.; and Yang, Y. 2018 · 2018
Cited alongside, same era.
nocaps: novel object captioning at scale
Agrawal, H.; Desai, K.; Wang, Y.; Chen, X.; Jain, R.; Johnson, M.; Batra, D.; Parikh, D.; Lee, S.; and Anderson, P. 2019 · 2019
Cited alongside, same era.
Wang, W.; Chen, Z.; and Hu, H. 2019 · 2019
Later among the works it cites.
Context and attribute grounded dense captioning
Yin, G.; Sheng, L.; Liu, B.; Yu, N.; Wang, X.; and Shao, J. 2019 · 2019
Later among the works it cites.
End-to-End Object Detection with Transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Closest in time.
UNITER: Learning universal image-text representations
Chen, Y.-C.; Li, L.; Yu, L.; Kholy, A. E.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
Closest in time.
Meshed-Memory Transformer for Image Captioning
Cornia, M.; Stefanini, M.; Baraldi, L.; and Cucchiara, R. 2020 · 2020
Closest in time.
Normalized and Geometry-Aware Self-Attention Network for Image Captioning
Guo, L.; Liu, J.; Zhu, X.; Yao, P.; Lu, S.; and Lu, H. 2020 · 2020
Closest in time.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Kuznetsova, A.; Rom, H.; Alldrin, N.; Uijlings, J.; Krasin, I.; Pont-Tuset, J.; Kamali, S.; Popov, S.; Malloci, M.; Duerig, T.; et al. 2020 · 2020
Closest in time.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X.; Yin, X.; Li, C.; Hu, X.; Zhang, P.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; Choi, Y.; and Gao, J. 2020 · 2020
Closest in time.
Learning to Generate Grounded Visual Captions without Localization Supervision
Ma, C.-Y.; Kalantidis, Y.; AlRegib, G.; Vajda, P.; Rohrbach, M.; and Kira, Z. 2020 · 2020
Closest in time.
X-Linear Attention Networks for Image Captioning
Pan, Y.; Yao, T.; Li, Y.; and Mei, T. 2020 · 2020
Closest in time.
TextCaps: a Dataset for Image Captioning with Reading Comprehension
Sidorov, O.; Hu, R.; Rohrbach, M.; and Singh, A. 2020 · 2020
Closest in time.
Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards
Yang, X.; Zhang, H.; Jin, D.; Liu, Y.; Wu, C.-H.; Tan, J.; Xie, D.; Wang, J.; and Wang, X. 2020 · 2020
Closest in time.