Fetching the paper…
Reading the bibliography…
In this paper we introduce the problem of Visual Semantic Role Labeling: given an image we want to detect people doing actions and localize the objects of interaction.
The berkeley framenet project
C. F. Baker, C. J. Fillmore, and J. B. Lowe · 1998
Earlier work this paper cites.
Recognizing human actions: a local svm approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Introduction to the conll-2005 shared task: Semantic role labeling
X. Carreras and L. Màrquez · 2005
Earlier work this paper cites.
VerbNet: A Broad-Coverage, Comprehensive Verb Lexicon
K. K. Schuler · 2006
Earlier work this paper cites.
Actions as space-time shapes
L. Gorelick, M. Blank, E. Shechtman, M. Irani, and R. Basri · 2007
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszałek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Action mach: a spatio-temporal maximum average correlation height filter for action recognition
M. Rodriguez, A. Javed, and M. Shah · 2008
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
A. Gupta, A. Kembhavi, and L. S. Davis · 2009
Earlier work this paper cites.
Understanding videos, constructing plots learning a visually grounded storyline model from annotated videos
A. Gupta, P. Srinivasan, J. Shi, and L. S. Davis · 2009
Earlier work this paper cites.
Actions in context
M. Marszałek, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
Discriminative subvolume search for efficient action detection
J. Yuan, Z. Liu, and Y. Wu · 2009
Earlier work this paper cites.
The PASCAL Visual Object Classes (VOC) Challenge
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Human action recognition by learning bases of action attributes and parts
B. Yao, X. Jiang, A. Khosla, A. L. Lin, L. Guibas, and L. Fei-Fei · 2011
Earlier work this paper cites.
Classifying actions and measuring action similarity by modeling the mutual context of objects and human poses
B. Yao, A. Khosla, and L. Fei-Fei · 2011
Cited alongside, same era.
Diagnosing error in object detectors
D. Hoiem, Y. Chodpathumwan, and Q. Dai · 2012
Cited alongside, same era.
Weakly supervised learning of interactions between humans and objects
A. Prest, C. Schmid, and V. Ferrari · 2012
Cited alongside, same era.
A database for fine grained activity detection of cooking activities
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele · 2012
Cited alongside, same era.
Recognizing human-object interactions in still images by modeling the mutual context of objects and human poses
B. Yao and L. Fei-Fei · 2012
Cited alongside, same era.
Towards understanding action recognition
H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black · 2013
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Later among the works it cites.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Later among the works it cites.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Later among the works it cites.
Microsoft COCO: Common objects in context
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Later among the works it cites.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2d human pose estimation: New benchmark and state of the art analysis
M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele · 2014
Cited alongside, same era.
Learning a recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
R-cnns for pose estimation and action detection
G. Gkioxari, B. Hariharan, R. Girshick, and J. Malik · 2014
Cited alongside, same era.
Finding action tubes
G. Gkioxari and J. Malik · 2014
Cited alongside, same era.
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Later among the works it cites.
Learning Deep Features for Scene Recognition using Places Database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Later among the works it cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. N. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2015
Closest in time.
R. Girshick · 2015
Closest in time.
Multiscale combinatorial grouping for image segmentation and object proposal generation
J. Pont-Tuset, P. Arbeláez, J. Barron, F. Marques, and J. Malik · 2015
Closest in time.
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Closest in time.