Fetching the paper…
Reading the bibliography…
The mainstream image captioning models rely on Convolutional Neural Network (CNN) image features to generate captions via recurrent models.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, H. Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Lei Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li · 2016
Earlier work this paper cites.
Review networks for caption generation
Zhilin Yang, Ye Yuan, Yuexin Wu, William W Cohen, and Russ R Salakhutdinov · 2016
Earlier work this paper cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher · 2017
Earlier work this paper cites.
Self-critical sequence training for image captioning
Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel · 2017
Earlier work this paper cites.
Scene graph generation by iterative message passing
Danfei Xu, Yuke Zhu, Christopher Choy, and Li Fei-Fei · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Earlier work this paper cites.
Women also snowboard: Overcoming bias in captioning models
Kaylee Burns, Lisa Anne Hendricks, Trevor Darrell, and Anna Rohrbach · 2018
Earlier work this paper cites.
Image captioning with scene-graph based semantic concepts
Lizhao Gao, Bo Wang, and Wenmin Wang · 2018
Earlier work this paper cites.
Factorizable net: An efficient subgraph-based framework for scene graph generation
Yikang Li, Wanli Ouyang, Bolei Zhou, Jianping Shi, Chao Zhang, and Xiaogang Wang · 2018
Cited alongside, same era.
Image captioning and visual question answering based on attributes and external knowledge
Qi Wu, Chunhua Shen, Peng Wang, A. Dick, and A. V. D. Hengel · 2018
Cited alongside, same era.
Scene graph captioner: Image captioning based on structural visual representation
Ning Xu, An-An Liu, Jing Liu, Weizhi Nie, and Yuting Su · 2018
Cited alongside, same era.
Graph r-cnn for scene graph generation
Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh · 2018
Cited alongside, same era.
Exploring visual relationship for image captioning
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei · 2018
Cited alongside, same era.
Neural motifs: Scene graph parsing with global context
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi · 2018
On the role of scene graphs in image captioning
Dalin Wang, Daniel Beck, and Trevor Cohn · 2019
Later among the works it cites.
Auto-encoding scene graphs for image captioning
Xu Yang, Kaihua Tang, Hanwang Zhang, and Jianfei Cai · 2019
Later among the works it cites.
Are scene graphs good enough to improve image captioning?
Victor Milewski, Marie-Francine Moens, and Iacer Calixto · 2020
Later among the works it cites.
A scene graph generation codebase in pytorch, 2020
Kaihua Tang · 2020
Later among the works it cites.
Unbiased scene graph generation from biased training
Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang · 2020
Later among the works it cites.
Vsgnet: Spatial attention network for detecting human object interactions using graph convolutions
Oytun Ulutan, A S M Iftekhar, and B. S. Manjunath · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unpaired image captioning via scene graph alignments
Jiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao, Xu Yang, and Gang Wang · 2019
Cited alongside, same era.
Scene graph generation with external knowledge and image reconstruction
Jiuxiang Gu, Handong Zhao, Zhe Lin, Sheng Li, Jianfei Cai, and Mingyang Ling · 2019
Cited alongside, same era.
Kuang-Huei Lee, Hamid Palangi, Xi Chen, Houdong Hu, and Jianfeng Gao · 2019
Cited alongside, same era.
Know more say less: Image captioning based on scene graphs
X. Li and S. Jiang · 2019
Cited alongside, same era.
Rethinking visual relationships for high-level image understanding, 02 2019
Yuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian, Li Zhu, and Tao Mei · 2019
Cited alongside, same era.
Learning to compose dynamic tree structures for visual contexts
Kaihua Tang, Hanwang Zhang, Baoyuan Wu, Wenhan Luo, and Wei Liu · 2019
Cited alongside, same era.
A survey of scene graph: Generation and application
Pengfei Xu, Xiaojun Chang, Ling Guo, Po-Yao Huang, Xiaojiang Chen, and Alex Hauptmann · 2020
Later among the works it cites.
PCPL: predicate-correlation perception learning for unbiased scene graph generation
Shaotian Yan, Chen Shen, Zhongming Jin, Jianqiang Huang, Rongxin Jiang, Yaowu Chen, and Xian-Sheng Hua · 2020
Later among the works it cites.
Sgae/ pytorch 0.4.0
Xu Yang · 2020
Later among the works it cites.
Comprehensive image captioning via scene graph decomposition
Yiwu Zhong, Liwei Wang, Jianshu Chen, Dong Yu, and Yin Li · 2020
Later among the works it cites.
Sub-gc/ pytorch 0.4.0
Yiwu Zhong · 2021
Closest in time.