Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Automatic image captioning
J.-Y. Pan, H.-J. Yang, P. Duygulu, and C. Faloutsos · 2004
Earlier work this paper cites.
The utility of affect expression in natural language interactions in joint human-robot tasks
M. Scheutz, P. Schermerhorn, and J. Kramer · 2006
Earlier work this paper cites.
Agency and communion from the perspective of self versus others
A. E. Abele and B. Wojciszke · 2007
Earlier work this paper cites.
Filling the emotion gap in linguistic theory: Commentary on potts’ expressive dimension
T. Jay and K. Janschewitz · 2007
Earlier work this paper cites.
The sixteen personality factor questionnaire (16pf)
H. E. Cattell and A. D. Mead · 2008
Earlier work this paper cites.
Torchvision the machine-vision package of torch
S. Marcel and Y. Rodriguez · 2010
Earlier work this paper cites.
What we instagram: A first analysis of instagram photo content and user types
Y. Hu, L. Manikonda, and S. Kambhampati · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Original
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Original
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Original
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
User conditional hashtag prediction for images
E. Denton, J. Weston, M. Paluri, L. Bourdev, and R. Fergus · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Original
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Multimodal convolutional neural networks for matching image and sentence
L. Ma, Z. Lu, L. Shang, and H. Li · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.