Fetching the paper…
Reading the bibliography…
Modeling instance-level context and object-object relationships is extremely challenging.
The effects of contextual scenes on the identification of objects
t. E. Palmer · 1975
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
First draft of a report on the edvac
J. Von Neumann · 1993
Earlier work this paper cites.
Gradient-based learning algorithms for recurrent networks and their computational complexity
R. J. Williams and D. Zipser · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Does consistent scene context facilitate object perception?
A. Hollingworth · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999
Earlier work this paper cites.
Using the forest to see the trees: a graphical model relating features, objects and scenes
K. Murphy, A. Torralba, W. Freeman, et al · 2003
Earlier work this paper cites.
Contextual priming for object detection
A. Torralba · 2003
Earlier work this paper cites.
Context-based vision system for place and object recognition
A. Torralba, K. P. Murphy, W. T. Freeman, M. A. Rubin, et al · 2003
Earlier work this paper cites.
Visual objects in context
M. Bar · 2004
Earlier work this paper cites.
A statistical model for general contextual object recognition
P. Carbonetto, N. De Freitas, and K. Barnard · 2004
Earlier work this paper cites.
Textonboost: Joint appearance, shape and context modeling for multi-class object recognition and segmentation
J. Shotton, J. Winn, C. Rother, and A. Criminisi · 2006
Earlier work this paper cites.
The role of context in object recognition
A. Oliva and A. Torralba · 2007
Earlier work this paper cites.
Objects in context
A. Rabinovich, A. Vedaldi, C. Galleguillos, E. Wiewiora, and S. Belongie · 2007
Earlier work this paper cites.
Object categorization using co-occurrence, location and appearance
C. Galleguillos, A. Rabinovich, and S. Belongie · 2008
Earlier work this paper cites.
Putting objects in perspective
D. Hoiem, A. A. Efros, and M. Hebert · 2008
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
An empirical study of context in object detection
S. K. Divvala, D. Hoiem, J. H. Hays, A. A. Efros, and M. Hebert · 2009
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
A. Gupta, A. Kembhavi, and L. S. Davis · 2009
Earlier work this paper cites.
Efficient subwindow search: A branch and bound framework for object localization
C. H. Lampert, M. B. Blaschko, and T. Hofmann · 2009
Earlier work this paper cites.
Beyond categories: The visual memex model for reasoning about object relationships
T. Malisiewicz and A. Efros · 2009
Earlier work this paper cites.
Actions in context
M. Marszalek, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Object detection with discriminatively trained part-based models
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan · 2010
Earlier work this paper cites.
Context based object categorization: A critical survey
C. Galleguillos and S. Belongie · 2010
Earlier work this paper cites.
Object bank: A high-level image representation for scene classification & semantic feature sparsification
L.-J. Li, H. Su, L. Fei-Fei, and E. P. Xing · 2010
Earlier work this paper cites.
Auto-context and its application to high-level vision tasks and 3d brain image segmentation
Z. Tu and X. Bai · 2010
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Discriminative models for multi-class object layout
C. Desai, D. Ramanan, and C. C. Fowlkes · 2011
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
P. Krähenbühl and V. Koltun · 2011
Earlier work this paper cites.
Extracting adaptive contextual cues from unlabeled regions
C. Li, D. Parikh, and T. Chen · 2011
Earlier work this paper cites.
Measuring the objectness of image windows
B. Alexe, T. Deselaers, and V. Ferrari · 2012
Cited alongside, same era.
Deepflow: Large displacement optical flow with deep matching
P. Weinzaepfel, J. Revaud, Z. Harchaoui, and C. Schmid · 2013
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
People watching: Human actions as a cue for single view geometry
D. F. Fouhey, V. Delaitre, A. Gupta, A. A. Efros, I. Laptev, and J. Sivic · 2014
Cited alongside, same era.
Hierarchical object detection with deep reinforcement learning
M. Bellver, X. Giró-i Nieto, F. Marqués, and J. Torres · 2016
Later among the works it cites.
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2016
Later among the works it cites.
Attend refine repeat: Active box proposal generation via in-out localization
S. Gidaris and N. Komodakis · 2016
Later among the works it cites.
Hybrid computing using a neural network with dynamic external memory
A. Graves, G. Wayne, M. Reynolds, T. Harley, I. Danihelka, A. Grabska-Barwińska, S. G. Colmenarejo, E. Grefenstette, T. Ramalho, J. Agapiou, et al · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2014
Cited alongside, same era.
The role of context for object detection and semantic segmentation in the wild
R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Cited alongside, same era.
Later among the works it cites.
End-to-end training of object class detectors for mean average precision
P. Henderson and V. Ferrari · 2016
Later among the works it cites.
A convnet for non-maximum suppression
J. Hosang, R. Benenson, and B. Schiele · 2016
Later among the works it cites.
Speed/accuracy trade-offs for modern convolutional object detectors
J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama, and K. Murphy · 2016
Later among the works it cites.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Later among the works it cites.
Attentive contexts for object detection
J. Li, Y. Wei, X. Liang, J. Dong, T. Xu, J. Feng, and S. Yan · 2016
Later among the works it cites.
R-fcn: Object detection via region-based fully convolutional networks
Y. Li, K. He, J. Sun, et al · 2016
Later among the works it cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2016
Later among the works it cites.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Later among the works it cites.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Later among the works it cites.
Adaptive object detection using adjacency and zoom prediction
Y. Lu, T. Javidi, and S. Lazebnik · 2016
Later among the works it cites.
Reinforcement learning for visual object detection
S. Mathe, A. Pirinen, and C. Sminchisescu · 2016
Later among the works it cites.
End-to-end instance segmentation and counting with recurrent attention
M. Ren and R. S. Zemel · 2016
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Later among the works it cites.
Where to look: Focus regions for visual question answering
K. J. Shih, S. Singh, and D. Hoiem · 2016
Later among the works it cites.
Contextual priming and feedback for faster r-cnn
A. Shrivastava and A. Gupta · 2016
Later among the works it cites.
Beyond Skip Connections: Top-Down Modulation for Object Detection
A. Shrivastava, R. Sukthankar, J. Malik, and A. Gupta · 2016
Later among the works it cites.
End-to-end people detection in crowded scenes
R. Stewart, M. Andriluka, and A. Y. Ng · 2016
Later among the works it cites.
Top-down learning for structured labeling with convolutional pseudoprior
S. Xie, X. Huang, and Z. Tu · 2016
Later among the works it cites.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Later among the works it cites.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
An implementation of faster rcnn with study for region sampling
X. Chen and A. Gupta · 2017
Closest in time.
Cognitive mapping and planning for visual navigation
S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik · 2017
Closest in time.
Deep variation-structured reinforcement learning for visual relationship and attribute detection
X. Liang, L. Lee, and E. P. Xing · 2017
Closest in time.
Neural map: Structured memory for deep reinforcement learning
E. Parisotto and R. Salakhutdinov · 2017
Closest in time.