The use of the area under the roc curve in the evaluation of machine learning algorithms
Andrew P Bradley · 1997
Earlier work this paper cites.
Interrater reliability: the kappa statistic
Mary L McHugh · 2012
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Original
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Exploring nearest neighbor approaches for image captioning
Original
Jacob Devlin, Saurabh Gupta, Ross Girshick, Margaret Mitchell, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Visual turing test for computer vision systems
Donald Geman, Stuart Geman, Neil Hallonquist, and Laurent Younes · 2015
Earlier work this paper cites.
Detection of cyberbullying incidents on the instagram social network
Original
Homa Hosseinmardi, Sabrina Arredondo Mattson, Rahat Ibn Rafiq, Richard Han, Qin Lv, and Shivakant Mishra · 2015
Earlier work this paper cites.
Expressing an image stream with a sequence of natural sentences
C. C. Park and G. Kim · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Original
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Recipe recognition with large multimodal food dataset
Xin Wang, Devinder Kumar, Nicolas Thome, Matthieu Cord, and Frederic Precioso · 2015
Earlier work this paper cites.
Simple baseline for visual question answering
Original
Bolei Zhou, Yuandong Tian, Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Original
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Grounding distributional semantics in the visual world
Marco Baroni · 2016
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Original
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visual storytelling
Ting-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, and Margaret Mitchell · 2016
Earlier work this paper cites.
Revisiting visual question answering baselines
Allan Jabri, Armand Joulin, and Laurens Van Der Maaten · 2016
Earlier work this paper cites.
Bag of tricks for efficient text classification
Original
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov · 2016
Earlier work this paper cites.
Virtual embodiment: A scalable long-term strategy for artificial intelligence research
Original
Douwe Kiela, Luana Bulat, Anita L Vero, and Stephen Clark · 2016
Earlier work this paper cites.
A shared task on multimodal machine translation and crosslingual image description
L. Specia, S. Frank, K. Sima’an, and D. Elliott · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Z. Waseem and D. Hovy · 2016
Earlier work this paper cites.
Are you a racist or am I seeing things? annotator influence on hate speech detection on twitter
Zeerak Waseem · 2016
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Original
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2016
Earlier work this paper cites.
Is a picture worth a thousand words? a deep multi-modal fusion architecture for product classification in e-commerce
Original
T. Zahavy, A. Magnani, A. Krishnan, and S. Mannor · 2016
Earlier work this paper cites.
Content-driven detection of cyberbullying on the instagram social network
Haoti Zhong, Hao Li, Anna Squicciarini, Sarah Rajtmajer, Christopher Griffin, David Miller, and Cornelia Caragea · 2016
Earlier work this paper cites.
Gated multimodal units for information fusion
John Arevalo, Thamar Solorio, Manuel Montes-y Gómez, and Fabio A González · 2017
Earlier work this paper cites.