Fetching the paper…
Reading the bibliography…
This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA).
Midwest and its children: The psychological ecology of an American town
R. Barker and H. Wright · 1954
Earlier work this paper cites.
The Hungarian method for the assignment problem
H. W. Kuhn · 1955
Earlier work this paper cites.
Word association norms, mutual information, and lexicoraphy
K.-W. Church and P. Hanks · 1990
Earlier work this paper cites.
Grammar of the film language
D. Arijon · 1991
Earlier work this paper cites.
Recognizing human actions: a local SVM approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Actions as space-time shapes
M. Blank, L. Gorelick, E. Shechtman, M. Irani, and R. Basri · 2005
Earlier work this paper cites.
Efficient visual event detection using volumetric features
Y. Ke, R. Sukthankar, and M. Hebert · 2005
Earlier work this paper cites.
Action MACH: a spatio-temporal maximum average correlation height filter for action recognition
M. Rodriguez, J. Ahmed, and M. Shah · 2008
Earlier work this paper cites.
Actions in context
M. Marszalek, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
Discriminative subvolume search for efficient action detection
J. Yuan, Z. Liu, and Y. Wu · 2009
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Towards understanding action recognition
H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. Black · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
TRECVID 2014 – an overview of the goals, tasks, data, evaluation mechanisms and metrics, 2014
P. Over, G. Awad, M. Michel, J. Fiscus, G. Sanders, W. Kraaij, A. Smeaton, and G. Quénot · 2014
Earlier work this paper cites.
ActivityNet: A large-scale video benchmark for human activity understanding
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles · 2015
Earlier work this paper cites.
HICO: A benchmark for recognizing human-object interactions in images
Y.-W. Chao, Z. Wang, Y. He, J. Wang, and J. Deng · 2015
Cited alongside, same era.
The PASCAL Visual Object Classes Challenge: A retrospective
M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2015
Cited alongside, same era.
Finding action tubes
G. Gkioxari and J. Malik · 2015
Cited alongside, same era.
S. Gupta and J. Malik · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Learning to track for spatio-temporal action localization
P. Weinzaepfel, Z. Harchaoui, and C. Schmid · 2015
Towards weakly-supervised action localization
P. Weinzaepfel, X. Martin, and C. Schmid · 2016
Later among the works it cites.
PersonNet: Person re-identification with deep convolutional neural networks
L. Wu, C. Shen, and A. van den Hengel · 2016
Later among the works it cites.
Quo vadis, action recognition? A new model and the Kinetics dataset
J. Carreira and A. Zisserman · 2017
Closest in time.
The “something something” video database for learning and evaluating visual common sense
R. Goyal, S. E. Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fründ, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic · 2017
Closest in time.
The devil is in the tails: Fine-grained classification in the wild
G. V. Horn and P. Perona · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
YouTube-8M: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Cited alongside, same era.
The Reel Truth: Women Aren’t Seen or Heard
Geena Davis Institute on Gender in Media · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Spot On: Action localization from pointly-supervised proposals
P. Mettes, J. van Gemert, and C. Snoek · 2016
Cited alongside, same era.
Multi-region two-stream R-CNN for action detection
X. Peng and C. Schmid · 2016
Cited alongside, same era.
Deep learning for detecting multiple space-time action tubes in videos
S. Saha, G. Singh, M. Sapienza, P. Torr, and F. Cuzzolin · 2016
Cited alongside, same era.
Closest in time.
Tube convolutional neural network (T-CNN) for action detection in videos
R. Hou, C. Chen, and M. Shah · 2017
Closest in time.
Speed/accuracy trade-offs for modern convolutional object detectors
J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama, and K. Murphy · 2017
Closest in time.
The THUMOS challenge on action recognition for videos “in the wild”
H. Idrees, A. R. Zamir, Y. Jiang, A. Gorban, I. Laptev, R. Sukthankar, and M. Shah · 2017
Closest in time.
FlowNet 2.0: Evolution of optical flow estimation with deep networks
E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox · 2017
Closest in time.
The Kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Closest in time.
AMTnet: Action-micro-tube regression by end-to-end trainable deep architecture
S. Saha, G.Sing, and F. Cuzzolin · 2017
Closest in time.
Online real-time multiple spatiotemporal action localisation and prediction
G. Singh, S. Saha, M. Sapienza, P. Torr, and F. Cuzzolin · 2017
Closest in time.
Action tubelet detector for spatio-temporal action localization
V.Kalogeiton, P. Weinzaepfel, V. Ferrari, and C. Schmid · 2017
Closest in time.
Every moment counts: Dense detailed labeling of actions in complex videos
S. Yeung, O. Russakovsky, N. Jin, M. Andriluka, G. Mori, and L. Fei-Fei · 2017
Closest in time.
SLAC: A sparsely labeled dataset for action classification and localization
H. Zhao, Z. Yan, H. Wang, L. Torresani, and A. Torralba · 2017
Closest in time.
M. Zolfaghari, G. Oliveira, N. Sedaghat, and T. Brox · 2017
Closest in time.