Fetching the paper…
Reading the bibliography…
The recent introduction of the AVA dataset for action detection has caused a renewed interest to this problem.
Wordnet: a lexical database for english
G. A. Miller · 1995
Earlier work this paper cites.
Data mining for direct marketing: Problems and solutions
C. X. Ling and C. Li · 1998
Earlier work this paper cites.
The case against accuracy estimation for comparing induction algorithms
F. J. Provost, T. Fawcett, R. Kohavi, et al · 1998
Earlier work this paper cites.
The class imbalance problem: A systematic study
N. Japkowicz and S. Stephen · 2002
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Human focused action localization in video
A. Kläser, M. Marszałek, C. Schmid, and A. Zisserman · 2010
Earlier work this paper cites.
Introduction to information retrieval
C. Manning, P. Raghavan, and H. Schütze · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
Explicit modeling of human-object interactions in realistic videos
A. Prest, V. Ferrari, and C. Schmid · 2013
Earlier work this paper cites.
Dense trajectories and motion boundary descriptors for action recognition
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2013
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Finding action tubes
G. Gkioxari and J. Malik · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Learning to track for spatio-temporal action localization
P. Weinzaepfel, Z. Harchaoui, and C. Schmid · 2015
Cited alongside, same era.
Multi-region two-stream r-cnn for action detection
X. Peng and C. Schmid · 2016
Cited alongside, same era.
Deep learning for detecting multiple space-time action tubes in videos
S. Saha, G. Singh, M. Sapienza, P. H. Torr, and F. Cuzzolin · 2016
Cited alongside, same era.
Action tubelet detector for spatio-temporal action localization
V. Kalogeiton, P. Weinzaepfel, V. Ferrari, and C. Schmid · 2017
Later among the works it cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. B. Girshick, K. He, and P. Dollár · 2017
Later among the works it cites.
Online real-time multiple spatiotemporal action localisation and prediction
G. Singh, S. Saha, M. Sapienza, P. H. Torr, and F. Cuzzolin · 2017
Later among the works it cites.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2017
Later among the works it cites.
Learning to model the tail
Y.-X. Wang, D. Ramanan, and M. Hebert · 2017
Later among the works it cites.
Open vocabulary scene parsing
H. Zhao, X. Puig, B. Zhou, S. Fidler, and A. Torralba · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training deep neural networks on imbalanced data sets
S. Wang, W. Liu, J. Wu, L. Cao, Q. Meng, and P. J. Kennedy · 2016
Cited alongside, same era.
Towards good practices for recognition & detection
Q. Zhong, C. Li, Y. Zhang, H. Sun, S. Yang, D. Xie, and S. Pu · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Deep pyramidal residual networks
D. Han, J. Kim, and J. Kim · 2017
Cited alongside, same era.
Mask R-CNN
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Cited alongside, same era.
Tube convolutional neural network (T-CNN) for action detection in videos
R. Hou, C. Chen, and M. Shah · 2017
Cited alongside, same era.
Large scale fine-grained categorization and domain-specific transfer learning
Y. Cui, Y. Song, C. Sun, A. Howard, and S. Belongie · 2018
Later among the works it cites.
Slowfast networks for video recognition
C. Feichtenhofer, H. Fan, J. Malik, and K. He · 2018
Later among the works it cites.
AVA: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, S. Vijayanarasimhan, C. Pantofaru, D. A. Ross, G. Toderici, Y. Li, S. Ricco, R. Sukthankar, C. Schmid, et al · 2018
Later among the works it cites.
D3D: Distilled 3d networks for video action recognition
J. C. Stroud, D. A. Ross, C. Sun, J. Deng, and R. Sukthankar · 2018
Later among the works it cites.
Video action transformer network
R. Girdhar, J. Carreira, C. Doersch, and A. Zisserman · 2019
Closest in time.
A structured model for action detection
Y. Zhang, P. Tokmakov, M. Hebert, and C. Schmid · 2019
Closest in time.