Fetching the paper…
Reading the bibliography…
Visual relationships capture a wide variety of interactions between pairs of objects in images (e.g.
Modeling the shape of the scene: A holistic representation of the spatial envelope
Oliva, A., Torralba, A.: · 2001
Earlier work this paper cites.
Dependency tree kernels for relation extraction
Culotta, A., Sorensen, J.: · 2004
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
Lowe, D.G.: · 2004
Earlier work this paper cites.
Exploring various knowledge in relation extraction
GuoDong, Z., Jian, S., Jie, Z., Min, Z.: · 2005
Earlier work this paper cites.
Discovering objects and their location in images
Sivic, J., Russell, B.C., Efros, A., Zisserman, A., Freeman, W.T., et al.: · 2005
Earlier work this paper cites.
Using multiple segmentations to discover objects and their extent in image collections
Russell, B.C., Freeman, W.T., Efros, A., Sivic, J., Zisserman, A., et al.: · 2006
Earlier work this paper cites.
Tree kernel-based relation extraction with context-sensitive structured parse tree information
ZHOU12, G., Zhang, M., Ji, D.H., Zhu, Q.: · 2007
Earlier work this paper cites.
Objects in context
Rabinovich, A., Vedaldi, A., Galleguillos, C., Wiewiora, E., Belongie, S.: · 2007
Earlier work this paper cites.
Towards scalable representations of object categories: Learning a hierarchy of parts
Fidler, S., Leonardis, A.: · 2007
Earlier work this paper cites.
Object categorization using co-occurrence, location and appearance
Galleguillos, C., Rabinovich, A., Belongie, S.: · 2008
Earlier work this paper cites.
Multi-class segmentation with relative location prior
Gould, S., Rodgers, J., Cohen, D., Elidan, G., Koller, D.: · 2008
Earlier work this paper cites.
Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers
Gupta, A., Davis, L.S.: · 2008
Earlier work this paper cites.
Putting objects in perspective
Hoiem, D., Efros, A.A., Hebert, M.: · 2008
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
Gupta, A., Kembhavi, A., Davis, L.S.: · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: · 2010
Earlier work this paper cites.
Graph cut based inference with co-occurrence statistics
Ladicky, L., Russell, C., Kohli, P., Torr, P.H.: · 2010
Earlier work this paper cites.
Context based object categorization: A critical survey
Galleguillos, C., Belongie, S.: · 2010
Cited alongside, same era.
Exploiting hierarchical context on a large database of object categories
Choi, M.J., Lim, J.J., Torralba, A., Willsky, A.S.: · 2010
Cited alongside, same era.
Modeling mutual context of object and human pose in human-object interaction activities
Yao, B., Fei-Fei, L.: · 2010
Cited alongside, same era.
Grouplet: A structured image representation for recognizing human and object interactions
Yao, B., Fei-Fei, L.: · 2010
Cited alongside, same era.
Efficiently selecting regions for scene understanding
Kumar, M.P., Koller, D.: · 2010
Cited alongside, same era.
Every picture tells a story: Generating sentences from images
Farhadi, A., Hejrati, M., Sadeghi, M.A., Young, P., Rashtchian, C., Hockenmaier, J., Forsyth, D.: · 2010
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
Guadarrama, S., Krishnamoorthy, N., Malkarnenkar, G., Venugopalan, S., Mooney, R., Darrell, T., Saenko, K.: · 2013
Later among the works it cites.
Grounding action descriptions in videos
Regneri, M., Rohrbach, M., Wetzel, D., Thater, S., Schiele, B., Pinkal, M.: · 2013
Later among the works it cites.
Learning the visual interpretation of sentences
Zitnick, C.L., Parikh, D., Vanderwende, L.: · 2013
Later among the works it cites.
Understanding indoor scenes using 3d geometric phrases
Choi, W., Chao, Y.W., Pantofaru, C., Savarese, S.: · 2013
Later among the works it cites.
Costa: Co-occurrence statistics for zero-shot classification
Mensink, T., Gavves, E., Snoek, C.G.: · 2014
Later among the works it cites.
Incorporating scene context and object layout into appearance modeling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Recognition using visual phrases
Sadeghi, M.A., Farhadi, A.: · 2011
Cited alongside, same era.
Learning to share visual appearance for multiclass object detection
Salakhutdinov, R., Torralba, A., Tenenbaum, J.: · 2011
Cited alongside, same era.
Action recognition from a distributed representation of pose and appearance
Maji, S., Bourdev, L., Malik, J.: · 2011
Cited alongside, same era.
Baby talk: Understanding and generating image descriptions
Kulkarni, G., Premraj, V., Dhar, S., Li, S., Choi, Y., Berg, A.C., Berg, T.L.: · 2011
Cited alongside, same era.
Semantic compositionality through recursive matrix-vector spaces
Socher, R., Huval, B., Manning, C.D., Ng, A.Y.: · 2012
Cited alongside, same era.
Describing the scene as a whole: Joint object detection, scene classification and semantic segmentation
Yao, J., Fidler, S., Urtasun, R.: · 2012
Cited alongside, same era.
Izadinia, H., Sadeghi, F., Farhadi, A.: · 2014
Later among the works it cites.
Integrating language and vision to generate natural language descriptions of videos in the wild
Thomason, J., Venugopalan, S., Guadarrama, S., Saenko, K., Mooney, R.: · 2014
Later among the works it cites.
From captions to visual concepts and back
Fang, H., Gupta, S., Iandola, F., Srivastava, R., Deng, L., Dollár, P., Gao, J., He, X., Mitchell, M., Platt, J., et al.: · 2014
Later among the works it cites.
Semantic parsing for text to 3d scene generation
Chang, A.X., Savva, M., Manning, C.D.: · 2014
Later among the works it cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., Malik, J.: · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K., Zisserman, A.: · 2014
Later among the works it cites.
Image retrieval using scene graphs
Johnson, J., Krishna, R., Stark, M., Li, L.J., Shamma, D.A., Bernstein, M., Fei-Fei, L.: · 2015
Later among the works it cites.
Learning semantic relationships for better action retrieval in images
Ramanathan, V., Li, C., Deng, J., Han, W., Li, Z., Gu, K., Song, Y., Bengio, S., Rossenberg, C., Fei-Fei, L.: · 2015
Later among the works it cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Schuster, S., Krishna, R., Chang, A., Fei-Fei, L., Manning, C.D.: · 2015
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., Bernstein, M., Fei-Fei, L.: · 2016
Closest in time.