Fetching the paper…
Reading the bibliography…
Despite progress in perceptual tasks such as image classification, computers still perform poorly on cognitive tasks such as image description and question answering.
80 million tiny images: A large data set for nonparametric object and scene recognition
Torralba, A., Fergus, R., and Freeman, W. T. (2008) · 1970
Earlier work this paper cites.
The naive physics manifesto
Hayes, P. J. (1978) · 1978
Earlier work this paper cites.
Qualitative process theory
Forbus, K. D. (1984) · 1984
Earlier work this paper cites.
The second naive physics manifesto
Hayes, P. J. (1985) · 1985
Earlier work this paper cites.
Culture and human development: A new look
Bruner, J. (1990) · 1990
Earlier work this paper cites.
Wordnet: a lexical database for english
Miller, G. A. (1995) · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
The berkeley framenet project
Baker, C. F., Fillmore, C. J., and Lowe, J. B. (1998) · 1998
Earlier work this paper cites.
Support vector machines
Hearst, M. A., Dumais, S. T., Osman, E., Platt, J., and Scholkopf, B. (1998) · 1998
Earlier work this paper cites.
Using corpus statistics and wordnet relations for sense identification
Leacock, C., Miller, G. A., and Chodorow, M. (1998) · 1998
Earlier work this paper cites.
A template-based approach toward acquisition of logical sentences
Hou, C.-S. J., Noy, N. F., and Musen, M. A. (2002) · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002) · 2002
Earlier work this paper cites.
Dependency tree kernels for relation extraction
Culotta, A. and Sorensen, J. (2004) · 2004
Earlier work this paper cites.
The senseval-3 english lexical sample task
Mihalcea, R., Chklovski, T. A., and Kilgarriff, A. (2004) · 2004
Earlier work this paper cites.
A shortest path dependency kernel for relation extraction
Bunescu, R. C. and Mooney, R. J. (2005) · 2005
Earlier work this paper cites.
Exploring various knowledge in relation extraction
GuoDong, Z., Jian, S., Jie, Z., and Min, Z. (2005) · 2005
Earlier work this paper cites.
Verbnet: A Broad-coverage, Comprehensive Verb Lexicon
Schuler, K. K. (2005) · 2005
Earlier work this paper cites.
A statistical approach to texture classification from single images
Varma, M. and Zisserman, A. (2005) · 2005
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Fei-Fei, L., Fergus, R., and Perona, P. (2007) · 2007
Earlier work this paper cites.
Learning visual attributes
Ferrari, V. and Zisserman, A. (2007) · 2007
Earlier work this paper cites.
Caltech-256 object category dataset
Griffin, G., Holub, A., and Perona, P. (2007) · 2007
Earlier work this paper cites.
Introduction to a large-scale general purpose ground truth database: methodology, annotation tool and benchmarks
Yao, B., Yang, X., and Zhu, S.-C. (2007) · 2007
Earlier work this paper cites.
Tree kernel-based relation extraction with context-sensitive structured parse tree information
Zhou, G., Zhang, M., Ji, D. H., and Zhu, Q. (2007) · 2007
Earlier work this paper cites.
Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers
Gupta, A. and Davis, L. S. (2008) · 2008
Earlier work this paper cites.
Labeled faces in the wild: A database forstudying face recognition in unconstrained environments
Huang, G. B., Mattar, M., Berg, T., and Learned-Miller, E. (2008) · 2008
Earlier work this paper cites.
Recognition by association via learning per-exemplar distances
Malisiewicz, T., Efros, A., et al. (2008) · 2008
Earlier work this paper cites.
Labelme: a database and web-based tool for image annotation
Russell, B. C., Torralba, A., Murphy, K. P., and Freeman, W. T. (2008) · 2008
Earlier work this paper cites.
Cheap and fast—but is it good?: evaluating non-expert annotations for natural language tasks
Snow, R., O’Connor, B., Jurafsky, D., and Ng, A. Y. (2008) · 2008
Earlier work this paper cites.
Toward never ending language learning
Betteridge, J., Carlson, A., Hong, S. A., Hruschka Jr, E. R., Law, E. L., Mitchell, T. M., and Wang, S. H. (2009) · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Describing objects by their attributes
Farhadi, A., Endres, I., Hoiem, D., and Forsyth, D. (2009) · 2009
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
Gupta, A., Kembhavi, A., and Davis, L. S. (2009) · 2009
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
Lampert, C. H., Nickisch, H., and Harmeling, S. (2009) · 2009
Earlier work this paper cites.
Statsnowball: a statistical approach to extracting entity relationships
Zhu, J., Nie, Z., Liu, X., Zhang, B., and Wen, J.-R. (2009) · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Everingham, M., Van Gool, L., Williams, C. K., Winn, J., and Zisserman, A. (2010) · 2010
Cited alongside, same era.
Every picture tells a story: Generating sentences from images
Farhadi, A., Hejrati, M., Sadeghi, M. A., Young, P., Rashtchian, C., Hockenmaier, J., and Forsyth, D. (2010) · 2010
Cited alongside, same era.
Building watson: An overview of the deepqa project
Ferrucci, D., Brown, E., Chu-Carroll, J., Fan, J., Gondek, D., Kalyanpur, A. A., Lally, A., Murdock, J. W., Nyberg, E., Prager, J., et al. (2010) · 2010
Cited alongside, same era.
Improving the fisher kernel for large-scale image classification
Perronnin, F., Sánchez, J., and Mensink, T. (2010) · 2010
Cited alongside, same era.
Sun database: Large-scale scene recognition from abbey to zoo
Xiao, J., Hays, J., Ehinger, K., Oliva, A., Torralba, A., et al. (2010) · 2010
Cited alongside, same era.
Modeling mutual context of object and human pose in human-object interaction activities
The sun attribute database: Beyond categories for deeper scene understanding
Patterson, G., Xu, C., Su, H., and Hays, J. (2014) · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Later among the works it cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2014) · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2014) · 2014
Later among the works it cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. (2014) · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yao, B. and Fei-Fei, L. (2010) · 2010
Cited alongside, same era.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T. L. (2011) · 2011
Cited alongside, same era.
Recognition using visual phrases
Sadeghi, M. A. and Farhadi, A. (2011) · 2011
Cited alongside, same era.
The caltech-ucsd birds-200-2011 dataset
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S. (2011) · 2011
Cited alongside, same era.
Pedestrian detection: An evaluation of the state of the art
Dollar, P., Wojek, C., Schiele, B., and Perona, P. (2012) · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Cited alongside, same era.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, P. K. and Fergus, R. (2012) · 2012
Cited alongside, same era.
Later among the works it cites.
Relation classification via convolutional deep neural network
Zeng, D., Liu, K., Lai, S., Zhou, G., and Zhao, J. (2014) · 2014
Later among the works it cites.
Reasoning about Object Affordances in a Knowledge Base Representation
Zhu, Y., Fathi, A., and Fei-Fei, L. (2014) · 2014
Later among the works it cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. (2015) · 2015
Later among the works it cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollar, P., and Zitnick, C. L. (2015) · 2015
Later among the works it cites.
Rmsprop and equilibrated adaptive learning rates for non-convex optimization
Dauphin, Y. N., de Vries, H., Chung, J., and Bengio, Y. (2015) · 2015
Later among the works it cites.
Cognition does not affect perception: Evaluating the evidence for “top-down” effects
Firestone, C. and Scholl, B. J. (2015) · 2015
Later among the works it cites.
Are you talking to a machine? dataset and methods for multilingual image question answering
Gao, H., Mao, J., Zhou, J., Huang, Z., Wang, L., and Xu, W. (2015) · 2015
Later among the works it cites.
Visual turing test for computer vision systems
Geman, D., Geman, S., Hallonquist, N., and Younes, L. (2015) · 2015
Later among the works it cites.
Girshick, R. (2015) · 2015
Later among the works it cites.
Discovering states and transformations in image collections
Isola, P., Lim, J. J., and Adelson, E. H. (2015) · 2015
Later among the works it cites.
Image retrieval using scene graphs
Johnson, J., Krishna, R., Stark, M., Li, L.-J., Shamma, D. A., Bernstein, M., and Fei-Fei, L. (2015) · 2015
Later among the works it cites.
Lebret, R., Pinheiro, P. O., and Collobert, R. (2015) · 2015
Later among the works it cites.
Learning to answer questions from image using convolutional neural network
Ma, L., Lu, Z., and Li, H. (2015) · 2015
Later among the works it cites.
Ask your neurons: A neural-based approach to answering questions about images
Malinowski, M., Rohrbach, M., and Fritz, M. (2015) · 2015
Later among the works it cites.
Word sense disambiguation: a survey
Pal, A. R. and Saha, D. (2015) · 2015
Later among the works it cites.
Learning semantic relationships for better action retrieval in images
Ramanathan, V., Li, C., Deng, J., Han, W., Li, Z., Gu, K., Song, Y., Bengio, S., Rossenberg, C., and Fei-Fei, L. (2015) · 2015
Later among the works it cites.
Autoextend: Extending word embeddings to embeddings for synsets and lexemes
Rothe, S. and Schütze, H. (2015) · 2015
Later among the works it cites.
Describing Common Human Visual Actions in Images
Ruggero Ronchi, M. and Perona, P. (2015) · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. (2015) · 2015
Later among the works it cites.
Viske: Visual knowledge extraction and question answering by visual verification of relation phrases
Sadeghi, F., Divvala, S. K., and Farhadi, A. (2015) · 2015
Later among the works it cites.
We are dynamo: Overcoming stalling and friction in collective action for crowd workers
Salehi, N., Irani, L. C., and Bernstein, M. S. (2015) · 2015
Later among the works it cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Schuster, S., Krishna, R., Chang, A., Fei-Fei, L., and Manning, C. D. (2015) · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A. C., Salakhutdinov, R., Zemel, R. S., and Bengio, Y. (2015) · 2015
Later among the works it cites.
Visual Madlibs: Fill in the blank Image Generation and Question Answering
Yu, L., Park, E., Berg, A. C., and Berg, T. L. (2015) · 2015
Later among the works it cites.
Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries
Zhu, Y., Zhang, C., Ré, C., and Fei-Fei, L. (2015) · 2015
Later among the works it cites.
Embracing error to enable rapid crowdsourcing
Krishna, R., Hata, K., Chen, S., Kravitz, J., Shamma, D. A., Fei-Fei, L., and Bernstein, M. S. (2016) · 2016
Closest in time.
Visual relationship detection using language priors
Lu, C., Krishna, R., Bernstein, M. S., and Fei-Fei, L. (2016) · 2016
Closest in time.
Yfcc100m: The new data in multimedia research
Thomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J. (2016) · 2016
Closest in time.