Fetching the paper…
Reading the bibliography…
Current state-of-the-art approaches for spatio-temporal action localization rely on detections at the frame level that are then linked or tracked across time.
High accuracy optical flow estimation based on a theory for warping
T. Brox, A. Bruhn, N. Papenberg, and J. Weickert · 2004
Earlier work this paper cites.
A survey on visual surveillance of object motion and behaviors
W. Hu, T. Tan, L. Wang, and S. Maybank · 2004
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results, 2007
M. Everingham, L. Van Gool, C. Williams, J. Winn, and A. Zisserman · 2007
Earlier work this paper cites.
Retrieving actions in movies
I. Laptev and P. Pérez · 2007
Earlier work this paper cites.
Action MACH: A spatio-temporal maximum average correlation height filter for action recognition
M. D. Rodriguez, J. Ahmed, and M. Shah · 2008
Earlier work this paper cites.
Cross-dataset action detection
L. Cao, Z. Liu, and T. S. Huang · 2010
Earlier work this paper cites.
Discriminative figure-centric models for joint action localization and recognition
T. Lan, Y. Wang, and G. Mori · 2011
Earlier work this paper cites.
A large-scale benchmark dataset for event recognition in surveillance video
S. Oh, A. Hoogs, A. Perera, N. Cuntoor, C.-C. Chen, J. T. Lee, S. Mukherjee, J. Aggarwal, H. Lee, L. Davis, et al · 2011
Earlier work this paper cites.
UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Towards understanding action recognition
H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black · 2013
Earlier work this paper cites.
Selective search for object recognition
J. R. Uijlings, K. E. van de Sande, T. Gevers, and A. W. Smeulders · 2013
Earlier work this paper cites.
Actionness ranking with lattice conditional ordinal random fields
W. Chen, C. Xiong, R. Xu, and J. Corso · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Action localization with tubelets from motion
M. Jain, J. Van Gemert, H. Jégou, P. Bouthemy, and C. G. Snoek · 2014
Cited alongside, same era.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Spatio-temporal object detection proposals
D. Oneata, J. Revaud, J. Verbeek, and C. Schmid · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Action detection by implicit intentional motion clustering
W. Chen and J. J. Corso · 2015
Cited alongside, same era.
APT: Action localization proposals from dense trajectories
J. Gemert, M. Jain, E. Gati, C. G. Snoek, et al · 2015
Cited alongside, same era.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Later among the works it cites.
Fast action proposals for human action detection and search
G. Yu and J. Yuan · 2015
Later among the works it cites.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Later among the works it cites.
Fast optical flow using dense inverse search
T. Kroeger, R. Timofte, D. Dai, and L. Van Gool · 2016
Later among the works it cites.
VideoLSTM convolves, attends and flows for action recognition
Z. Li, E. Gavves, M. Jain, and C. G. Snoek · 2016
Later among the works it cites.
SSD: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C. Y. Fu, and A. C. Berg · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finding action tubes
G. Gkioxari and J. Malik · 2015
Cited alongside, same era.
Unsupervised tube extraction using transductive learning and dense trajectories
M. Marian Puscas, E. Sangineto, D. Culibrk, and N. Sebe · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Sequence to sequence—video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Cited alongside, same era.
Learning to track for spatio-temporal action localization
P. Weinzaepfel, Z. Harchaoui, and C. Schmid · 2015
Cited alongside, same era.
Multi-region two-stream R-CNN for action detection
X. Peng and C. Schmid · 2016
Later among the works it cites.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2016
Later among the works it cites.
Deep learning for detecting multiple space-time action tubes in videos
S. Saha, G. Singh, M. Sapienza, P. H. Torr, and F. Cuzzolin · 2016
Later among the works it cites.
Actionness estimation using hybrid fully convolutional networks
L. Wang, Y. Qiao, X. Tang, and L. Van Gool · 2016
Later among the works it cites.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Later among the works it cites.
Online real time multiple spatiotemporal action localisation and prediction on a single platform
G. Singh, S. Saha, M. Sapienza, P. Torr, and F. Cuzzolin · 2017
Closest in time.