Fetching the paper…
Reading the bibliography…
We aim to tackle a novel vision task called Weakly Supervised Visual Relation Detection (WSVRD) to detect "subject-predicate-object" relations in an image with object relation groundtruths available only at the image level.
A framework for multiple-instance learning
O. Maron and T. Lozano-Pérez · 1998
Earlier work this paper cites.
Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers
A. Gupta and L. S. Davis · 2008
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
A. Gupta, A. Kembhavi, and L. S. Davis · 2009
Earlier work this paper cites.
Self-paced learning for latent variable models
M. P. Kumar, B. Packer, and D. Koller · 2010
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Recognition using visual phrases
M. A. Sadeghi and A. Farhadi · 2011
Earlier work this paper cites.
Detecting actions, poses, and objects with relational phraselets
C. Desai and D. Ramanan · 2012
Earlier work this paper cites.
Weakly supervised learning of interactions between humans and objects
A. Prest, C. Schmid, and V. Ferrari · 2012
Earlier work this paper cites.
Restoring an image taken through a window covered with dirt or rain
D. Eigen, D. Krishnan, and R. Fergus · 2013
Earlier work this paper cites.
Selective search for object recognition
J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
On learning to localize objects with minimal supervision
H. O. Song, R. B. Girshick, S. Jegelka, J. Mairal, Z. Harchaoui, T. Darrell, et al · 2014
Earlier work this paper cites.
Weakly supervised object localization with latent category learning
C. Wang, W. Ren, K. Huang, and T. Tan · 2014
Earlier work this paper cites.
Edge boxes: Locating object proposals from edges
C. L. Zitnick and P. Dollár · 2014
Earlier work this paper cites.
Leveraging linguistic structure for open domain information extraction
G. Angeli, M. J. Premkumar, and C. D. Manning · 2015
Earlier work this paper cites.
Hico: A benchmark for recognizing human-object interactions in images
Y.-W. Chao, Z. Wang, Y. He, J. Wang, and J. Deng · 2015
Earlier work this paper cites.
Fast r-cnn
R. Girshick · 2015
Cited alongside, same era.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Cited alongside, same era.
Learning semantic relationships for better action retrieval in images
V. Ramanathan, C. Li, J. Deng, W. Han, Z. Li, K. Gu, Y. Song, S. Bengio, C. Rossenberg, and L. Fei-Fei · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
S. Schuster, R. Krishna, A. Chang, L. Fei-Fei, and C. D. Manning · 2015
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Later among the works it cites.
Yfcc100m: The new data in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2016
Later among the works it cites.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Q. Wu, P. Wang, C. Shen, A. Dick, and A. van den Hengel · 2016
Later among the works it cites.
Situation recognition: Visual semantic role labeling for image understanding
M. Yatskar, L. Zettlemoyer, and A. Farhadi · 2016
Later among the works it cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Learning to generalize to new compositions in image understanding
Y. Atzmon, J. Berant, V. Kezami, A. Globerson, and G. Chechik · 2016
Cited alongside, same era.
Weakly supervised deep detection networks
H. Bilen and A. Vedaldi · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2016
Cited alongside, same era.
Later among the works it cites.
Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning
L. Chen, H. Zhang, J. Xiao, L. Nie, J. Shao, W. Liu, and T.-S. Chua · 2017
Closest in time.
Weakly supervised object localization with multi-fold multiple instance learning
R. G. Cinbis, J. Verbeek, and C. Schmid · 2017
Closest in time.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Closest in time.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2017
Closest in time.
Deep self-taught learning for weakly supervised object localization
Z. Jie, Y. Wei, X. Jin, J. Feng, and W. Liu · 2017
Closest in time.
Vip-cnn: Visual phrase guided convolutional neural network
Y. Li, W. Ouyang, and X. Wang · 2017
Closest in time.
Scene graph generation from objects, phrases and region captions
Y. Li, W. Ouyang, B. Zhou, K. Wang, and X. Wang · 2017
Closest in time.
Surveillance video parsing with single frame supervision
S. Liu, C. Wang, R. Qian, H. Yu, R. Bao, and Y. Sun · 2017
Closest in time.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
Closest in time.
Relationship proposal networks
J. Zhang, M. Elhoseiny, S. Cohen, W. Chang, and A. Elgammal · 2017
Closest in time.