Seeing through the human reporting bias: Visual classifiers from noisy human-centric labels
I. Misra, C. Lawrence Zitnick, M. Mitchell, and R. Girshick · 2016
Later among the works it cites.
Generating natural questions about an image
N. Mostafazadeh, I. Misra, J. Devlin, M. Mitchell, X. He, and L. Vanderwende · 2016
Later among the works it cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Later among the works it cites.
Learning to ask: Neural question generation for reading comprehension
Original
X. Du, J. Shao, and C. Cardie · 2017
Later among the works it cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Later among the works it cites.
Learning visual n-grams from web data
A. Li, A. Jabri, A. Joulin, and L. van der Maaten · 2017
Later among the works it cites.
Self-critical sequence training for image captioning
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel · 2017
Later among the works it cites.
Context-aware captions from context-agnostic supervision
R. Vedantam, S. Bengio, K. Murphy, D. Parikh, and G. Chechik · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Later among the works it cites.
http://lsun.cs.princeton.edu/slides/caption_open.pdf
Microsoft COCO 1st Captioning Challenge (Large-scale Scene UNderstanding Workshop, CVPR 2015) · 2018
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Later among the works it cites.
Convolutional image captioning
J. Aneja, A. Deshpande, and A. G. Schwing · 2018
Later among the works it cites.
Vse++: Improving visual-semantic embeddings with hard negatives
F. Faghri, D. J. Fleet, J. R. Kiros, and S. Fidler · 2018
Later among the works it cites.
Stacked cross attention for image-text matching
K.-H. Lee, X. Chen, G. Hua, H. Hu, and X. He · 2018
Later among the works it cites.
Discriminability objective for training descriptive captions
R. Luo, B. Price, S. Cohen, and G. Shakhnarovich · 2018
Later among the works it cites.
Exploring the limits of weakly supervised pretraining
D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten · 2018
Later among the works it cites.
Learning by asking questions
I. Misra, R. Girshick, R. Fergus, M. Hebert, A. Gupta, and L. van der Maaten · 2018
Later among the works it cites.
Engaging image captioning via personality
Original
K. Shuster, S. Humeau, H. Hu, A. Bordes, and J. Weston · 2018
Later among the works it cites.
A corpus for reasoning about natural language grounded in photographs
A. Suhr, S. Zhou, I. Zhang, H. Bai, and Y. Artzi · 2018
Later among the works it cites.