Fetching the paper…
Reading the bibliography…
Image captioning models generally lack the capability to take into account user interest, and usually default to global descriptions that try to balance readability, informativeness, and information overload.
Graph-rise: Graph-regularized image semantic embedding
Da-Cheng Juan, Chun-Ta Lu, Zhen Li, Futang Peng, Aleksei Timofeev, Yi-Ting Chen, Yaxi Gao, Tom Duerig, Andrew Tomkins, and Sujith Ravi. 2019 · 1902
Earlier work this paper cites.
Captioning transformer with stacked attention modules
Xinxin Zhu, Lixiang Li, Jing Liu, Haipeng Peng, and Xinxin Niu. 2018 · 2012
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
Girish Kulkarni, Visruth Premraj, Vicente Ordonez, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg. 2013 · 2013
Earlier work this paper cites.
Multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Rich Zemel. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Mind’s eye: A recurrent visual representation for image caption generation
Xinlei Chen and Lawrence C Zitnick. 2015 · 2015
Earlier work this paper cites.
Language models for image captioning: The quirks and what works
Jacob Devlin, Hao Cheng, Hao Fang, Saurabh Gupta, Li Deng, Xiaodong He, Geoffrey Zweig, and Margaret Mitchell. 2015 · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Earlier work this paper cites.
Describing images using inferred visual dependency representations
Desmond Elliott and Arjen de Vries. 2015 · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li. 2015 · 2015
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (mRNN)
J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille. 2015 · 2015
Cited alongside, same era.
Image captioning with an intermediate attributes layer
Qi Wu, Chunhua Shen, Anton van den Hengel, Lingqiao Liu, and Anthony Dick. 2015 · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei. 2016 · 2016
Cited alongside, same era.
Image captioning with semantic attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo. 2016 · 2016
Cited alongside, same era.
Neural baby talk
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2018 · 2018
Later among the works it cites.
Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Later among the works it cites.
Decoupled box proposal and featurization with ultrafine-grained semantic labels improve image captioning and visual question answering
Soravit Changpinyo, Bo Pang, Piyush Sharma, and Radu Soricut. 2019 · 2019
Later among the works it cites.
Show, control and tell: a framework for generating controllable and grounded captions
Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2019 · 2019
Later among the works it cites.
Fast, diverse and accurate image captioning guided by part-of-speech
Aditya Deshpande, Jyoti Aneja, Liwei Wang, Alexander G Schwing, and David Forsyth. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guided open vocabulary image captioning with constrained beam search
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2017 · 2017
Cited alongside, same era.
A hierarchical approach for generating descriptive image paragraphs
Jonathan Krause, Justin Johnson, Ranjay Krishna, and Li Fei-Fei. 2017 · 2017
Cited alongside, same era.
Text-guided attention model for image captioning
Jonghwan Mun, Minsu Cho, and Bohyung Han. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
A dataset for Sanskrit word segmentation
Amrith Krishna, Pavan Kumar Satuluri, and Pawan Goyal. 2017a
Cited in the paper.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael Bernstein, and Li Fei-Fei. 2017b
Cited in the paper.
Intention oriented image captions with guiding objects
Yue Zheng, Yali Li, and Shengjin Wang. 2019 · 2019
Later among the works it cites.
Say as you wish: Fine-grained control of image caption generation with abstract scene graphs
Shizhe Chen, Qin Jin, Peng Wang, and Qi Wu. 2020 · 2020
Closest in time.
The Open Images Dataset V4
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper R. R. Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Tom Duerig, and Vittorio Ferrari. 2020 · 2020
Closest in time.
Connecting vision and language with localized narratives
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, and Vittorio Ferrari. 2020 · 2020
Closest in time.
Unified vision-language pre-training for image captioning and VQA
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2020 · 2020
Closest in time.