Fetching the paper…
Reading the bibliography…
Descriptive region features extracted by object detection networks have played an important role in the recent advancements of image captioning.
BLEU: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S.; and Lavie, A. 2005 · 2005
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O.; Toshev, A.; Bengio, S.; and Erhan, D. 2015 · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A.; Salakhudinov, R.; Zemel, R.; and Bengio, Y. 2015 · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Anderson, P.; Fernando, B.; Johnson, M.; and Gould, S. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Lu, J.; Xiong, C.; Parikh, D.; and Socher, R. 2017 · 2017
Cited alongside, same era.
Self-critical sequence training for image captioning
Rennie, S. J.; Marcheret, E.; Mroueh, Y.; Ross, J.; and Goel, V. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Aggregated Residual Transformations for Deep Neural Networks
Xie, S.; Girshick, R. B.; Dollár, P.; Tu, Z.; and He, K. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Exploring visual relationship for image captioning
Yao, T.; Pan, Y.; Li, Y.; and Mei, T. 2018 · 2018
Hierarchy Parsing for Image Captioning
Yao, T.; Pan, Y.; Li, Y.; and Mei, T. 2019 · 2019
Later among the works it cites.
Meshed-Memory Transformer for Image Captioning
Cornia, M.; Stefanini, M.; Baraldi, L.; and Cucchiara, R. 2020 · 2020
Later among the works it cites.
Normalized and Geometry-Aware Self-Attention Network for Image Captioning
Guo, L.; Liu, J.; Zhu, X.; Yao, P.; Lu, S.; and Lu, H. 2020 · 2020
Later among the works it cites.
In Defense of Grid Features for Visual Question Answering
Jiang, H.; Misra, I.; Rohrbach, M.; Learned-Miller, E.; and Chen, X. 2020 · 2020
Later among the works it cites.
X-Linear Attention Networks for Image Captioning
Pan, Y.; Yao, T.; Li, Y.; and Mei, T. 2020 · 2020
Later among the works it cites.
Reinforcing an Image Caption Generator Using Off-Line Human Feedback
Seo, P. H.; Sharma, P.; Levinboim, T.; Han, B.; and Soricut, R. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Image Captioning: Transforming Objects into Words
Herdade, S.; Kappeler, A.; Boakye, K.; and Soares, J. 2019 · 2019
Cited alongside, same era.
Attention on Attention for Image Captioning
Huang, L.; Wang, W.; Chen, J.; and Wei, X.-Y. 2019 · 2019
Cited alongside, same era.
Hierarchical attention network for image captioning
Wang, W.; Chen, Z.; and Hu, H. 2019 · 2019
Cited alongside, same era.
Auto-encoding scene graphs for image captioning
Yang, X.; Tang, K.; Zhang, H.; and Cai, J. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Show, Recall, and Tell: Image Captioning with Recall Mechanism
Wang, L.; Bai, Z.; Zhang, Y.; and Lu, H. 2020 · 2020
Later among the works it cites.
MemCap: Memorizing Style Knowledge for Image Captioning
Zhao, W.; Wu, X.; and Zhang, X. 2020 · 2020
Later among the works it cites.
Unified Vision-Language Pre-Training for Image Captioning and VQA
Zhou, L.; Palangi, H.; Zhang, L.; Hu, H.; Corso, J. J.; and Gao, J. 2020 · 2020
Later among the works it cites.
Squeeze-and-Excitation Networks
Hu, J.; Shen, L.; Albanie, S.; Sun, G.; and Wu, E. 2020 · 2023
Closest in time.