Fetching the paper…
Reading the bibliography…
Most scene graph parsers use a two-stage pipeline to detect visual relationships: the first stage detects entities, and the second predicts the predicate for each entity pair using a softmax distribution.
Towards context-aware interaction recognition for visual relationship detection
B. Zhuang, L. Liu, C. Shen, and I. Reid · 1903
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Learning word embeddings efficiently with noise-contrastive estimation
A. Mnih and K. Kavukcuoglu · 2013
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, R. S. Zemel, and et al · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. Plummer, L. Wang, C. Cervantes, J. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Earlier work this paper cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Earlier work this paper cites.
Order-embeddings of images and language
I. Vendrov, R. Kiros, S. Fidler, and R. Urtasun · 2016
Earlier work this paper cites.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Query-guided regression network with context policy for phrase grounding
K. Chen, R. Kovvuri, and R. Nevatia · 2017
Earlier work this paper cites.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Cited alongside, same era.
Aligned image-word representations improve inductive transfer across vision-language tasks
T. Gupta, K. J. Shih, S. Singh, and D. Hoiem · 2017
Cited alongside, same era.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2017
Cited alongside, same era.
Openimages: A public dataset for large-scale multi-label and multi-class image classification
I. Krasin, T. Duerig, N. Alldrin, V. Ferrari, S. Abu-El-Haija, A. Kuznetsova, H. Rom, J. Uijlings, S. Popov, S. Kamali, M. Malloci, J. Pont-Tuset, A. Veit, S. Belongie, V. Gomes, A. Gupta, C. Sun, G. Chechik, D. Cai, Z. Feng, D. Narayanan, and K. Murphy · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Cited alongside, same era.
Visual relationship detection with internal and external linguistic knowledge distillation
R. Yu, A. Li, V. I. Morariu, and L. S. Davis · 2017
Later among the works it cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
S. Zagoruyko and N. Komodakis · 2017
Later among the works it cites.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
Later among the works it cites.
Ppr-fcn: Weakly supervised visual relation detection via parallel pairwise r-fcn
H. Zhang, Z. Kyaw, J. Yu, and S.-F. Chang · 2017
Later among the works it cites.
Relationship proposal networks
J. Zhang, M. Elhoseiny, S. Cohen, W. Chang, and A. Elgammal · 2017
Later among the works it cites.
Detecting and recognizing human-object intaractions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vip-cnn: A visual phrase reasoning convolutional neural network for visual relationship detection
Y. Li, W. Ouyang, and X. Wang · 2017
Cited alongside, same era.
Deep variation-structured reinforcement learning for visual relationship and attribute detection
X. Liang, L. Lee, and E. P. Xing · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie · 2017
Cited alongside, same era.
Referring expression generation and comprehension via attributes
J. Liu, L. Wang, M.-H. Yang, et al · 2017
Cited alongside, same era.
Comprehension-guided referring expressions
R. Luo and G. Shakhnarovich · 2017
Cited alongside, same era.
Pixels to graphs by associative embedding
A. Newell and J. Deng · 2017
Cited alongside, same era.
Weakly-supervised learning of visual relations
J. Peyre, I. Laptev, C. Schmid, and J. Sivic · 2017
Cited alongside, same era.
G. Gkioxari, R. Girshick, P. Dollár, and K. He · 2018
Later among the works it cites.
Graph r-cnn for scene graph generation
J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh · 2018
Later among the works it cites.
Shuffle-then-assemble: Learning object-agnostic visual relationship features
X. Yang, H. Zhang, and J. Cai · 2018
Later among the works it cites.
Zoom-net: Mining deep feature interactions for visual relationship recognition
G. Yin, L. Sheng, B. Liu, N. Yu, X. Wang, J. Shao, and C. Change Loy · 2018
Later among the works it cites.
Mattnet: Modular attention network for referring expression comprehension
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg · 2018
Later among the works it cites.
Neural motifs: Scene graph parsing with global context
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi · 2018
Later among the works it cites.
An interpretable model for scene graph generation
J. Zhang, K. Shih, A. Tao, B. Catanzaro, and A. Elgammal · 2018
Later among the works it cites.
Introduction to the 1st place winning model of openimages relationship detection challenge
J. Zhang, K. Shih, A. Tao, B. Catanzaro, and A. Elgammal · 2018
Later among the works it cites.
Large-scale visual relationship understanding
J. Zhang, Y. Kalantidis, M. Rohrbach, M. Paluri, A. Elgammal, and M. Elhoseiny · 2019
Closest in time.