Fetching the paper…
Reading the bibliography…
We introduce a simple yet surprisingly powerful model to incorporate attention in action recognition and human object interaction tasks.
Visual routines
S. Ullman · 1984
Earlier work this paper cites.
Automatic annotation of everyday movements
D. Ramanan and D. A. Forsyth · 2003
Earlier work this paper cites.
Is bottom-up attention useful for object recognition?
U. Rutishauser, D. Walther, C. Koch, and P. Perona · 2004
Earlier work this paper cites.
An integrated model of top-down and bottom-up attention for optimizing detection speed
V. Navalpakkam and L. Itti · 2006
Earlier work this paper cites.
Fisher kernels on visual vocabularies for image categorization
F. Perronnin and C. Dance · 2007
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
A. Gupta, A. Kembhavi, and L. S. Davis · 2009
Earlier work this paper cites.
Recognizing human actions in still images: a study of bag-of-features and part-based representations
V. Delaitre, I. Laptev, and J. Sivic · 2010
Earlier work this paper cites.
Discriminative models for static human-object interactions
C. Desai, D. Ramanan, and C. Fowlkes · 2010
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Recognizing human actions from still images with latent poses
W. Yang, Y. Wang, and G. Mori · 2010
Earlier work this paper cites.
Grouplet: A structured image representation for recognizing human and object interactions
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Mechanisms of top-down attention
F. Baluch and L. Itti · 2011
Earlier work this paper cites.
Learning person-object interactions for action recognition in still images
V. Delaitre, J. Sivic, and I. Laptev · 2011
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Action recognition from a distributed representation of pose and appearance
S. Maji, L. Bourdev, and J. Malik · 2011
Earlier work this paper cites.
Action Recognition by Dense Trajectories
H. Wang, A. Kläser, C. Schmid, and L. Cheng-Lin · 2011
Earlier work this paper cites.
Human action recognition by learning bases of action attributes and parts
B. Yao, X. Jiang, A. Khosla, A. Lin, L. Guibas, and L. Fei-Fei · 2011
Earlier work this paper cites.
Semantic segmentation with second-order pooling
J. Carreira, R. Caseiro, J. Batista, and C. Sminchisescu · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
THUMOS challenge: Action recognition with a large number of classes
Y. Jiang, J. Liu, A. Roshan Zamir, I. Laptev, M. Piccardi, M. Shah, and R. Sukthankar · 2013
Cited alongside, same era.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Cited alongside, same era.
Fine-grained activity recognition with holistic and pose based features
L. Pishchulin, M. Andriluka, and B. Schiele · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Hico: A benchmark for recognizing human-object interactions in images
Y.-W. Chao, Z. Wang, Y. He, J. Wang, and J. Deng · 2015
Cited alongside, same era.
P-CNN: Pose-based CNN Features for Action Recognition
G. Chéron, I. Laptev, and C. Schmid · 2015
Cited alongside, same era.
Action recognition using visual attention
S. Sharma, R. Kiros, and R. Salakhutdinov · 2016
Later among the works it cites.
Joint network based attention for action recognition
Y. Shi, Y. Tian, Y. Wang, and T. Huang · 2016
Later among the works it cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, and V. Vanhoucke · 2016
Later among the works it cites.
Long-term temporal convolutions for action recognition
G. Varol, I. Laptev, and C. Schmid · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, S. V. M. Rohrbach, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Contextual action recognition with R*CNN
G. Gkioxari, R. Girshick, and J. Malik · 2015
Cited alongside, same era.
SALICON: Reducing the semantic gap in saliency prediction by adapting deep neural networks
X. Huang, C. Shen, X. Boix, and Q. Zhao · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Bilinear CNN models for fine-grained visual recognition
T.-Y. Lin, A. RoyChowdhury, and S. Maji · 2015
Cited alongside, same era.
Describing common human visual actions in images
M. Ronchi and P. Perona · 2015
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Later among the works it cites.
Learning deep features for discriminative localization
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2016
Later among the works it cites.
Realtime multi-person 2d pose estimation using part affinity fields
Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh · 2017
Closest in time.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Closest in time.
ActionVLAD: Learning spatio-temporal aggregation for action classification
R. Girdhar, D. Ramanan, A. Gupta, J. Sivic, and B. Russell · 2017
Closest in time.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Closest in time.
Hadamard Product for Low-rank Bilinear Pooling
J.-H. Kim, K. W. On, W. Lim, J. Kim, J.-W. Ha, and B.-T. Zhang · 2017
Closest in time.
Low-rank bilinear pooling for fine-grained classification
S. Kong and C. Fowlkes · 2017
Closest in time.
Learning latent sub-events in activity videos using temporal attention filters
A. Piergiovanni, C. Fan, and M. S. Ryoo · 2017
Closest in time.
A simple neural network module for relational reasoning
A. Santoro, D. Raposo, D. G. Barrett, M. Malinowski, R. Pascanu, P. Battaglia, and T. Lillicrap · 2017
Closest in time.
An end-to-end spatio-temporal attention model for human action recognition from skeleton data
S. Song, C. Lan, J. Xing, W. Zeng, and J. Liu · 2017
Closest in time.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Closest in time.
Residual attention network for image classification
F. Wang, M. Jiang, C. Qian, S. Yang, C. Li, H. Zhang, X. Wang, and X. Tang · 2017
Closest in time.
Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection
M. Zolfaghari, G. L. Oliveira, N. Sedaghat, and T. Brox · 2017
Closest in time.