Fetching the paper…
Reading the bibliography…
We propose a new model for detecting visual relationships, such as "person riding motorcycle" or "bottle on table".
Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers
A. Gupta and L. S. Davis · 2008
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
A. Gupta, A. Kembhavi, and L. S. Davis · 2009
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Recognition using visual phrases
M. A. Sadeghi and A. Farhadi · 2011
Earlier work this paper cites.
Detecting actions, poses, and objects with relational phraselets
C. Desai and D. Ramanan · 2012
Earlier work this paper cites.
Weakly supervised learning of interactions between humans and objects
A. Prest, C. Schmid, and V. Ferrari · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2015
Earlier work this paper cites.
S. Gupta and J. Malik · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
W. Chan, N. Jaitly, Q. Le, and O. Vinyals · 2016
Earlier work this paper cites.
Chained predictions using convolutional neural networks
G. Gkioxari, A. Toshev, and N. Jaitly · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Pixel recurrent neural networks
A. van den Oord, N. Kalchbrenner, and K. Kavukcuoglu · 2016
Focal loss for dense object detection
T. Lin, P. Goyal, R. B. Girshick, K. He, and P. Dollár · 2017
Later among the works it cites.
Weakly-supervised learning of visual relations
J. Peyre, I. Laptev, C. Schmid, and J. Sivic · 2017
Later among the works it cites.
PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications
T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Later among the works it cites.
Visual relationship detection with internal and external linguistic knowledge distillation
R. Yu, A. Li, V. I. Morariu, and L. S. Davis · 2017
Later among the works it cites.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Conditional image generation with PixelCNN decoders
A. van den Oord, N. Kalchbrenner, O. Vinyals, L. Espeholt, A. Graves, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Cited alongside, same era.
Speed/accuracy trade-offs for modern convolutional object detectors
J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama, et al · 2017
Cited alongside, same era.
Video pixel networks
N. Kalchbrenner, A. van den Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu · 2017
Cited alongside, same era.
ViP-CNN: Visual phrase guided convolutional neural network
Y. Li, W. Ouyang, X. Wang, and X. Tang · 2017
Cited alongside, same era.
Deep variation-structured reinforcement learning for visual relationship and attribute detection
X. Liang, L. Lee, and E. P. Xing · 2017
Cited alongside, same era.
Later among the works it cites.
PPR-FCN: weakly supervised visual relation detection via parallel pairwise R-FCN
H. Zhang, Z. Kyaw, J. Yu, and S.-F. Chang · 2017
Later among the works it cites.
ican: Instance-centric attention network for human-object interaction detection
C. Gao, Y. Zou, and J.-B. Huang · 2018
Closest in time.
Detecting and recognizing human-object interactions
G. Gkioxari, R. Girshick, P. Dollár, and K. He · 2018
Closest in time.
Referring relationships
R. Krishna, I. Chami, M. S. Bernstein, and L. Fei-Fei · 2018
Closest in time.
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, T. Duerig, and V. Ferrari · 2018
Closest in time.
Neural motifs: Scene graph parsing with global context
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi · 2018
Closest in time.