Fetching the paper…
Reading the bibliography…
Rich semantic relations are important in a variety of visual recognition problems.
Global training of document processing systems using graph transformer networks
L. Bottou, Y. Bengio, and Y. Le Cun · 1997
Earlier work this paper cites.
Context-based vision system for place and object recognition
A. Torralba, K. P. Murphy, W. T. Freeman, and M. A. Rubin · 2003
Earlier work this paper cites.
Putting objects in perspective
D. Hoiem, A. Efros, and M. Hebert · 2006
Earlier work this paper cites.
What are they doing? : Collective activity classification using spatio-temporal relationship among people
W. Choi, K. Shahid, and S. Savarese · 2009
Earlier work this paper cites.
Understanding videos, constructing plots learning a visually grounded storyline model from annotated videos
A. Gupta, P. Srinivasan, J. Shi, and L. S. Davis · 2009
Earlier work this paper cites.
Beyond actions: Discriminative models for contextual group activities
T. Lan, Y. Wang, W. Yang, and G. Mori · 2010
Earlier work this paper cites.
Learning context for collective activity recognition
W. Choi, K. Shahid, and S. Savarese · 2011
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
P. Krähenbühl and V. Koltun · 2011
Earlier work this paper cites.
Learning message-passing inference machines for structured prediction
S. Ross, D. Munoz, M. Hebert, and J. A. Bagnell · 2011
Earlier work this paper cites.
Stochastic representation and recognition of high-level group activities
M. Ryoo and J. Aggarwal · 2011
Earlier work this paper cites.
Cost-sensitive top-down/bottom-up inference for multiscale activity recognition
M. R. Amer, D. Xie, M. Zhao, S. Todorovic, and S.-C. Zhu · 2012
Earlier work this paper cites.
A unified framework for multi-target tracking and collective activity recognition
W. Choi and S. Savarese · 2012
Earlier work this paper cites.
Combining per-frame and per-track cues for multi-person action recognition
S. Khamis, V. I. Morariu, and L. S. Davis · 2012
Earlier work this paper cites.
A flow model for joint action recognition and identity maintenance
S. Khamis, V. I. Morariu, and L. S. Davis · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Social roles in hierarchical models for human activity recognition
T. Lan, L. Sigal, and G. Mori · 2012
Cited alongside, same era.
Discriminative latent models for recognizing contextual group activities
T. Lan, Y. Wang, W. Yang, S. Robinovitch, and G. Mori · 2012
Cited alongside, same era.
Multi-agent event detection: Localization and role assignment
S. Kwak, B. Han, and J. H. Han · 2013
Cited alongside, same era.
Context-aware modeling and recognition of activities in video
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Later among the works it cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Later among the works it cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Joint training of a convolutional network and a graphical model for human pose estimation
J. J. Tompson, A. Jain, Y. LeCun, and C. Bregler · 2014
Later among the works it cites.
Deep structured models for group activity recognition
Z. Deng, M. Zhai, L. Chen, Y. Liu, S. Muralidharan, M. Roshtkhari, , and G. Mori · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Zhu, N. M. Nayak, and A. K. Roy-Chowdhury · 2013
Cited alongside, same era.
Hirf: Hierarchical random field for collective activity recognition in videos
M. R. Amer, P. Lei, and S. Todorovic · 2014
Cited alongside, same era.
Hirf: Hierarchical random field for collective activity recognition in videos
M. R. Amer, P. Lei, and S. Todorovic · 2014
Cited alongside, same era.
Learning latent constituents for recognition of group activities in video
B. Antic and B. Ommer · 2014
Cited alongside, same era.
Semantic image segmentation with deep convolutional nets and fully connected crfs
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2014
Cited alongside, same era.
Learning deep structured models
L.-C. Chen, A. G. Schwing, A. L. Yuille, and R. Urtasun · 2014
Cited alongside, same era.
Discovering groups of people in images
W. Choi, Y. Chao, C. Pantofaru, and S. Savarese · 2014
Cited alongside, same era.
Probabilistic label relation graphs with ising models
N. Ding, J. Deng, K. Murphy, and H. Neven · 2015
Closest in time.
Learning ensembles of potential functions for structured prediction with latent variables
H. Hajimirsadeghi and G. Mori · 2015
Closest in time.
Visual recognition by counting instances: A multi-instance cardinality potential kernel
H. Hajimirsadeghi, W. Yan, A. Vahdat, and G. Mori · 2015
Closest in time.
Fully connected deep structured networks
A. G. Schwing and R. Urtasun · 2015
Closest in time.
Joint inference of groups, events and human roles in aerial videos
T. Shu, D. Xie, B. Rothrock, S. Todorovic, and S.-C. Zhu · 2015
Closest in time.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Closest in time.
Improving object detection with deep convolutional networks via bayesian optimization and structured prediction
Y. Zhang, K. Sohn, R. Villegas, G. Pan, and H. Lee · 2015
Closest in time.
Conditional random fields as recurrent neural networks
S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. S. Torr · 2015
Closest in time.