Fetching the paper…
Reading the bibliography…
Existing image captioning systems are dedicated to generating narrative captions for images, which are spatially detached from the image in presentation.
Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Meteor Universal: Language Specific Translation Evaluation for Any Target Language. In Proceedings of the Ninth Workshop on Statistical Machine Translation . 376–380
Michael J. Denkowski and Alon Lavie. 2014 · 2014
Earlier work this paper cites.
Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. In IEEE Conference on Computer Vision and Pattern Recognition . 580–587
Ross B. Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In European Conference on Computer Vision . 740–755
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks. In International Conference on Learning Representations
Pierre Sermanet, David Eigen, Xiang Zhang, Michaël Mathieu, Rob Fergus, and Yann LeCun. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Annual Conference on Neural Information Processing Systems . 91–99
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition . 4566–4575
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator. In IEEE Conference on Computer Vision and Pattern Recognition . 3156–3164
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision and Pattern Recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
DenseCap: Fully Convolutional Localization Networks for Dense Captioning. In IEEE Conference on Computer Vision and Pattern Recognition . 4565–4574
Justin Johnson, Andrej Karpathy, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Rico: A Mobile App Dataset for Building Data-Driven Design Applications. In Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology . 845–854
Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017b · 2017
Earlier work this paper cites.
Towards End-to-End Text Spotting with Convolutional Recurrent Neural Networks. In IEEE International Conference on Computer Vision . 5248–5256
Hui Li, Peng Wang, and Chunhua Shen. 2017 · 2017
Earlier work this paper cites.
Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training. In IEEE International Conference on Computer Vision . 4155–4164
Rakshith Shetty, Marcus Rohrbach, Lisa Anne Hendricks, Mario Fritz, and Bernt Schiele. 2017 · 2017
Earlier work this paper cites.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In IEEE Conference on Computer Vision and Pattern Recognition . 6077–6086
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Earlier work this paper cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna M. Wallach, Hal Daumé III, and Kate Crawford. 2018 · 2018
Cited alongside, same era.
FOTS: Fast Oriented Text Spotting With a Unified Network. In IEEE Conference on Computer Vision and Pattern Recognition . 5676–5685
Xuebo Liu, Ding Liang, Shi Yan, Dagui Chen, Yu Qiao, and Junjie Yan. 2018 · 2018
Cited alongside, same era.
Entity-aware Image Caption Generation. In Conference on Empirical Methods in Natural Language Processing . 4013–4023
Di Lu, Spencer Whitehead, Lifu Huang, Heng Ji, and Shih-Fu Chang. 2018 · 2018
Cited alongside, same era.
Training for Diversity in Image Paragraph Captioning. In Conference on Empirical Methods in Natural Language Processing . 757–761
Luke Melas-Kyriazi, Alexander M. Rush, and George Han. 2018 · 2018
Cited alongside, same era.
BreakingNews: Article Annotation by Image and Text Processing
Arnau Ramisa, Fei Yan, Francesc Moreno-Noguer, and Krystian Mikolajczyk. 2018 · 2018
UNITER: UNiversal Image-TExt Representation Learning. In European Conference on Computer Vision . 104–120
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Meshed-memory transformer for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10578–10587
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. 2020 · 2020
Later among the works it cites.
Layout Generation and Completion with Self-attention
Kamal Gupta, Alessandro Achille, Justin Lazarow, Larry Davis, Vijay Mahadevan, and Abhinav Shrivastava. 2020 · 2020
Later among the works it cites.
Captioning Images Taken by People Who Are Blind. In European Conference on Computer Vision . 417–434
Danna Gurari, Yinan Zhao, Meng Zhang, and Nilavra Bhattacharya. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Good News, Everyone! Context Driven Entity-Aware Captioning for News Images. In IEEE Conference on Computer Vision and Pattern Recognition . 12466–12475
Ali Furkan Biten, Lluís Gómez, Marçal Rusiñol, and Dimosthenis Karatzas. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Conference of the North American Chapter of the Association for Computational Linguistics . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Unified Language Model Pre-training for Natural Language Understanding and Generation. In Annual Conference on Neural Information Processing Systems . 13042–13054
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 2019
Cited alongside, same era.
Attention on Attention for Image Captioning. In IEEE International Conference on Computer Vision . 4633–4642
Lun Huang, Wenmin Wang, Jie Chen, and Xiaoyong Wei. 2019 · 2019
Cited alongside, same era.
LayoutVAE: Stochastic Scene Layout Generation From a Label Set. In IEEE International Conference on Computer Vision . 9894–9903
Akash Abdu Jyothi, Thibaut Durand, Jiawei He, Leonid Sigal, and Greg Mori. 2019 · 2019
Cited alongside, same era.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In Annual Conference on Neural Information Processing Systems . 13–23
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, and Daxin Jiang. 2020 · 2020
Later among the works it cites.
X-Linear Attention Networks for Image Captioning. In IEEE Conference on Computer Vision and Pattern Recognition . 10968–10977
Yingwei Pan, Ting Yao, Yehao Li, and Tao Mei. 2020 · 2020
Later among the works it cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
TextCaps: A Dataset for Image Captioning with Reading Comprehension. In European Conference on Computer Vision . 742–758
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh. 2020 · 2020
Later among the works it cites.
Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards. In European Conference on Computer Vision . 1–17
Xuewen Yang, Heming Zhang, Di Jin, Yingru Liu, Chi-Hao Wu, Jianchao Tan, Dongliang Xie, Jue Wang, and Xin Wang. 2020 · 2020
Later among the works it cites.
Variational Transformer Networks for Layout Generation. In IEEE Conference on Computer Vision and Pattern Recognition . 13642–13652
Diego Martín Arroyo, Janis Postels, and Federico Tombari. 2021 · 2021
Later among the works it cites.
Towards Diverse Paragraph Captioning for Untrimmed Videos. In IEEE Conference on Computer Vision and Pattern Recognition . 11245–11254
Yuqing Song, Shizhe Chen, and Qin Jin. 2021 · 2021
Later among the works it cites.
End-to-End Dense Video Captioning With Parallel Decoding. In IEEE International Conference on Computer Vision . 6847–6857
Teng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng, Ran Cheng, and Ping Luo. 2021 · 2021
Later among the works it cites.
TAP: Text-Aware Pre-Training for Text-VQA and Text-Caption. In IEEE Conference on Computer Vision and Pattern Recognition . 8751–8761
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florêncio, Lijuan Wang, Cha Zhang, Lei Zhang, and Jiebo Luo. 2021 · 2021
Later among the works it cites.
RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual Words. In IEEE Conference on Computer Vision and Pattern Recognition . 15465–15474
Xuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji, Yiyi Zhou, Yongjian Wu, Feiyue Huang, and Rongrong Ji. 2021 · 2021
Later among the works it cites.
Kaleido-BERT: Vision-Language Pre-Training on Fashion Domain. In IEEE Conference on Computer Vision and Pattern Recognition . 12647–12657
Mingchen Zhuge, Dehong Gao, Deng-Ping Fan, Linbo Jin, Ben Chen, Haoming Zhou, Minghui Qiu, and Ling Shao. 2021 · 2021
Later among the works it cites.