Fetching the paper…
Reading the bibliography…
This paper introduces a novel approach for modeling visual relations between pairs of objects.
Diffrac: a discriminative and flexible framework for clustering
F. R. Bach and Z. Harchaoui · 2008
Earlier work this paper cites.
Object categorization using co-occurrence, location and appearance
C. Galleguillos, A. Rabinovich, and S. Belongie · 2008
Earlier work this paper cites.
Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers
A. Gupta and L. S. Davis · 2008
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
A. Gupta, A. Kembhavi, and L. S. Davis · 2009
Earlier work this paper cites.
Discriminative models for static human-object interactions
C. Desai, D. Ramanan, and C. Fowlkes · 2010
Earlier work this paper cites.
Grouplet: A structured image representation for recognizing human and object interactions
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Learning person-object interactions for action recognition in still images
V. Delaitre, J. Sivic, and I. Laptev · 2011
Earlier work this paper cites.
Weakly supervised learning of interactions between humans and objects
A. Prest, C. Schmid, and V. Ferrari · 2011
Earlier work this paper cites.
Recognition using visual phrases
M. A. Sadeghi and A. Farhadi · 2011
Earlier work this paper cites.
Human action recognition by learning bases of action attributes and parts
B. Yao, X. Jiang, A. Khosla, A. L. Lin, L. Guibas, and L. Fei-Fei · 2011
Earlier work this paper cites.
Context models and out-of-context objects
M. J. Choi, A. Torralba, and A. S. Willsky · 2012
Earlier work this paper cites.
A latent factor model for highly multi-relational data
R. Jenatton, N. L. Roux, A. Bordes, and G. R. Obozinski · 2012
Earlier work this paper cites.
Automatic discovery of groups of objects for scene understanding
C. Li, D. Parikh, and T. Chen · 2012
Earlier work this paper cites.
Neil: Extracting visual knowledge from web data
X. Chen, A. Shrivastava, and A. Gupta · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. A. Ranzato, and T. Mikolov · 2013
Earlier work this paper cites.
Reasoning with neural tensor networks for knowledge base completion
R. Socher, D. Chen, C. D. Manning, and A. Ng · 2013
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
R. Socher, M. Ganjoo, C. D. Manning, and A. Ng · 2013
Earlier work this paper cites.
Weakly supervised action labeling in videos under ordering constraints
P. Bojanowski, R. Lajugie, F. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
Efficient image and video co-localization with frank-wolfe algorithm
A. Joulin, K. Tang, and L. Fei-Fei · 2014
Cited alongside, same era.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and L. Fei-Fei · 2014
Cited alongside, same era.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Cited alongside, same era.
Is this a wampimuk? cross-modal mapping between distributional semantics and the visual world
Viske: Visual knowledge extraction and question answering by visual verification of relation phrases
F. Sadeghi, S. K. Divvala, and A. Farhadi · 2015
Later among the works it cites.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Later among the works it cites.
Weakly supervised deep detection networks
H. Bilen and A. Vedaldi · 2016
Later among the works it cites.
Sherlock: Scalable fact learning in images
M. Elhoseiny, S. Cohen, W. Chang, B. Price, and A. Elgammal · 2016
Later among the works it cites.
Deep compositional captioning: Describing novel object categories without paired training data
L. A. Hendricks, S. Venugopalan, M. Rohrbach, R. Mooney, K. Saenko, and T. Darrell · 2016
Later among the works it cites.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Lazaridou, E. Bruni, and M. Baroni · 2014
Cited alongside, same era.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Reasoning about object affordances in a knowledge base representation
Y. Zhu, A. Fathi, and L. Fei-Fei · 2014
Cited alongside, same era.
Text to 3d scene generation with rich lexical grounding
A. Chang, W. Monroe, M. Savva, C. Potts, and C. D. Manning · 2015
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. N. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2015
Cited alongside, same era.
Fast R-CNN
R. Girshick · 2015
Cited alongside, same era.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei · 2015
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Later among the works it cites.
Minding the gaps for block Frank-Wolfe optimization of structured SVMs
A. Osokin, J.-B. Alayrac, I. Lukasewitz, P. K. Dokania, and S. Lacoste-Julien · 2016
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Later among the works it cites.
Captioning images with diverse objects
S. Venugopalan, L. A. Hendricks, M. Rohrbach, R. Mooney, T. Darrell, and K. Saenko · 2016
Later among the works it cites.
Latent embeddings for zero-shot classification
Y. Xian, Z. Akata, G. Sharma, Q. Nguyen, M. Hein, and B. Schiele · 2016
Later among the works it cites.
Stating the obvious: Extracting visual common sense knowledge
M. Yatskar, V. Ordonez, and A. Farhadi · 2016
Later among the works it cites.
A multipath network for object detection
S. Zagoruyko, A. Lerer, T.-Y. Lin, P. O. Pinheiro, S. Gross, S. Chintala, and P. Dollár · 2016
Later among the works it cites.
Learning from video and text via large-scale discriminative clustering
A. Miech, J.-B. Alayrac, P. Bojanowski, I. Laptev, and J. Sivic · 2017
Closest in time.