Fetching the paper…
Reading the bibliography…
Visual attention not only improves the performance of image captioners, but also serves as a visual interpretation to qualitatively measure the caption rationality and model transparency.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
Siming Li, Girish Kulkarni, Tamara L Berg, Alexander C Berg, and Yejin Choi · 2011
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
Margaret Mitchell, Xufeng Han, Jesse Dodge, Alyssa Mensch, Amit Goyal, Alex Berg, Kota Yamaguchi, Tamara Berg, Karl Stratos, and Hal Daumé III · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
Girish Kulkarni, Visruth Premraj, Vicente Ordonez, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg · 2013
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Cross modal distillation for supervision transfer
Saurabh Gupta, Judy Hoffman, and Jitendra Malik · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele · 2016
Cited alongside, same era.
Learning deep structure-preserving image-text embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2016
Cited alongside, same era.
Image captioning with semantic attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo · 2016
Cited alongside, same era.
Query-guided regression network with context policy for phrase grounding
Kan Chen, Rama Kovvuri, and Ram Nevatia · 2017
Cited alongside, same era.
Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning
Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, Wei Liu, and Tat-Seng Chua · 2017
Cited alongside, same era.
An empirical study of language cnn for image captioning
Jiuxiang Gu, Gang Wang, Jianfei Cai, and Tsuhan Chen · 2017
Multi-label image classification via knowledge distillation from weakly-supervised detection
Yongcheng Liu, Lu Sheng, Jing Shao, Junjie Yan, Shiming Xiang, and Chunhong Pan · 2018
Later among the works it cites.
Neural baby talk
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2018
Later among the works it cites.
Discriminability objective for training descriptive captions
Ruotian Luo, Brian Price, Scott Cohen, and Gregory Shakhnarovich · 2018
Later among the works it cites.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko · 2018
Later among the works it cites.
Exploring visual relationship for image captioning
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei · 2018
Later among the works it cites.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Modeling relationships in referential expressions with compositional modular networks
Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Cited alongside, same era.
Attention correctness in neural image captioning
Chenxi Liu, Junhua Mao, Fei Sha, and Alan Yuille · 2017
Cited alongside, same era.
Referring expression generation and comprehension via attributes
Jingyu Liu, Liang Wang, and Ming-Hsuan Yang · 2017
Cited alongside, same era.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher · 2017
Cited alongside, same era.
Hierarchical multimodal lstm for dense visual-semantic embedding
Zhenxing Niu, Mo Zhou, Le Wang, Xinbo Gao, and Gang Hua · 2017
Cited alongside, same era.
Later among the works it cites.
Align2ground: Weakly supervised phrase grounding guided by image-caption alignment
Samyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja, Devi Parikh, and Ajay Divakaran · 2019
Later among the works it cites.
Aligning linguistic words and visual semantic units for image captioning
Longteng Guo, Jing Liu, Jinhui Tang, Jiangwei Li, Wei Luo, and Hanqing Lu · 2019
Later among the works it cites.
Attention on attention for image captioning
Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei · 2019
Later among the works it cites.
Learning to assemble neural module tree networks for visual grounding
Daqing Liu, Hanwang Zhang, Zheng-Jun Zha, and Wu Feng · 2019
Later among the works it cites.
Referring expression grounding by marginalizing scene graph likelihood
Daqing Liu, Hanwang Zhang, Zheng-Jun Zha, and Fanglin Wang · 2019
Later among the works it cites.
Learning to generate grounded image captions without localization supervision
Chih-Yao Ma, Yannis Kalantidis, Ghassan AlRegib, Peter Vajda, Marcus Rohrbach, and Zsolt Kira · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton · 2019
Later among the works it cites.
Matching images and text with multi-modal tensor fusion and re-ranking
Tan Wang, Xing Xu, Yang Yang, Alan Hanjalic, Heng Tao Shen, and Jingkuan Song · 2019
Later among the works it cites.
Auto-encoding scene graphs for image captioning
Xu Yang, Kaihua Tang, Hanwang Zhang, and Jianfei Cai · 2019
Later among the works it cites.
Learning to collocate neural modules for image captioning
Xu Yang, Hanwang Zhang, and Jianfei Cai · 2019
Later among the works it cites.
Ckd: Cross-task knowledge distillation for text-to-image synthesis
Mingkuan Yuan and Yuxin Peng · 2019
Later among the works it cites.
Context-aware visual policy network for fine-grained image captioning
Zheng-Jun Zha, Daqing Liu, Hanwang Zhang, Yongdong Zhang, and Feng Wu · 2019
Later among the works it cites.
Grounded video description
Luowei Zhou, Yannis Kalantidis, Xinlei Chen, Jason J Corso, and Marcus Rohrbach · 2019
Later among the works it cites.
Unbiased scene graph generation from biased training
Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang · 2020
Closest in time.
Deconfounded image captioning: A causal retrospect
Xu Yang, Hanwang Zhang, and Jianfei Cai · 2020
Closest in time.