Fetching the paper…
Reading the bibliography…
As the intermediate level task connecting image captioning and object detection, visual relationship detection started to catch researchers' attention because of its descriptive power and clear structure.
Generalization of backpropagation with application to a recurrent gas market model
P. J. Werbos · 1988
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Matching words and pictures
K. Barnard, P. Duygulu, D. Forsyth, N. d. Freitas, D. M. Blei, and M. I. Jordan · 2003
Earlier work this paper cites.
Using multiple segmentations to discover objects and their extent in image collections
B. C. Russell, W. T. Freeman, A. A. Efros, J. Sivic, and A. Zisserman · 2006
Earlier work this paper cites.
Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers
A. Gupta and L. S. Davis · 2008
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Efficiently selecting regions for scene understanding
M. P. Kumar and D. Koller · 2010
Earlier work this paper cites.
Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora
R. Socher and L. Fei-Fei · 2010
Earlier work this paper cites.
Learning cross-modality similarity for multinomial data
Y. Jia, M. Salzmann, and T. Darrell · 2011
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
V. Koltun · 2011
Earlier work this paper cites.
Recognition using visual phrases
M. A. Sadeghi and A. Farhadi · 2011
Earlier work this paper cites.
Measuring the objectness of image windows
B. Alexe, T. Deselaers, and V. Ferrari · 2012
Earlier work this paper cites.
Detecting actions, poses, and objects with relational phraselets
C. Desai and D. Ramanan · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2013
Earlier work this paper cites.
Generalizing image captions for image-text parallel corpus
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Learning a recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
Spatial pyramid pooling in deep convolutional networks for visual recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2015
Later among the works it cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, et al · 2015
Cited alongside, same era.
Fast r-cnn
R. Girshick · 2015
Cited alongside, same era.
Deformable part models are convolutional neural networks
R. Girshick, F. Iandola, T. Darrell, and J. Malik · 2015
Cited alongside, same era.
Later among the works it cites.
Conditional random fields as recurrent neural networks
S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. Torr · 2015
Later among the works it cites.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2015
Later among the works it cites.
Learning to generalize to new compositions in image understanding
Y. Atzmon, J. Berant, V. Kezami, A. Globerson, and G. Chechik · 2016
Later among the works it cites.
Structured feature learning for pose estimation
X. Chu, W. Ouyang, H. Li, and X. Wang · 2016
Later among the works it cites.
Crf-cnn: Modeling structured information in human pose estimation
X. Chu, W. Ouyang, X. Wang, et al · 2016
Later among the works it cites.
T-cnn: Tubelets with convolutional neural networks for object detection from videos
K. Kang, H. Li, J. Yan, X. Zeng, B. Yang, T. Xiao, C. Zhang, Z. Wang, R. Wang, X. Wang, et al · 2016
Later among the works it cites.
Object detection from video tubelets with convolutional neural networks
K. Kang, W. Ouyang, H. Li, and X. Wang · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Later among the works it cites.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Dense captioning with joint inference and visual context
L. Yang, K. Tang, J. Yang, and L.-J. Li · 2016
Later among the works it cites.
Image captioning with semantic attention
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Later among the works it cites.
Object detection in videos with tubelet proposal networks
K. Kang, H. Li, T. Xiao, W. Ouyang, J. Yan, X. Liu, and X. Wang · 2017
Closest in time.