Fetching the paper…
Reading the bibliography…
We present SlowFast networks for video recognition.
Receptive fields and functional architecture in two non-striate visual areas of the cat
D. H. Hubel and T. N. Wrisel · 1965
Earlier work this paper cites.
Spatial and temporal contrast sensitivities of neurones in lateral geniculate nucleus of macaque
A. Derrington and P. Lennie · 1984
Earlier work this paper cites.
Spatiotemporal energy models for the perception of motion
E. H. Adelson and J. R. Bergen · 1985
Earlier work this paper cites.
Segregation of form, color, movement, and depth: anatomy, physiology, and perception
M. Livingstone and D. Hubel · 1988
Earlier work this paper cites.
Distributed hierarchical processing in the primate cerebral cortex
D. J. Felleman and D. C. Van Essen · 1991
Earlier work this paper cites.
The statistics of natural images
D. L. Ruderman · 1994
Earlier work this paper cites.
Neural mechanisms of form and motion processing in the primate visual system
D. C. Van Essen and J. L. Gallant · 1994
Earlier work this paper cites.
Statistics of natural images and models
J. Huang and D. Mumford · 1999
Earlier work this paper cites.
Motion illusions as optimal percepts
Y. Weiss, E. P. Simoncelli, and E. H. Adelson · 2002
Earlier work this paper cites.
Behavior recognition via sparse spatio-temporal features
P. Dollár, V. Rabaud, G. Cottrell, and S. Belongie · 2005
Earlier work this paper cites.
Human detection using oriented histograms of flow and appearance
N. Dalal, B. Triggs, and C. Schmid · 2006
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
A. Kläser, M. Marszałek, and C. Schmid · 2008
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Convolutional learning of spatio-temporal features
G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler · 2010
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Fast R-CNN
R. Girshick · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3D convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
A short note about Kinetics-600
J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman · 2018
Closest in time.
Spatio-temporal channel correlation networks for action classification
A. Diba, M. Fayyaz, V. Sharma, M. M. Arzani, R. Yousefzadeh, J. Gall, and L. Van Gool · 2018
Closest in time.
The ActivityNet large-scale activity recognition challenge 2018 summary
B. Ghanem, J. C. Niebles, C. Snoek, F. C. Heilbron, H. Alwassel, V. Escorcia, R. Khrisna, S. Buch, and C. D. Dao · 2018
Closest in time.
R. Girdhar, J. Carreira, C. Doersch, and A. Zisserman · 2018
Closest in time.
Detectron
R. Girshick, I. Radosavovic, G. Gkioxari, P. Dollár, and K. He · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Spatiotemporal residual networks for video action recognition
C. Feichtenhofer, A. Pinz, and R. Wildes · 2016
Cited alongside, same era.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Val Gool · 2016
Cited alongside, same era.
AVA: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik · 2018
Closest in time.
Exploiting spatial-temporal modelling and multi-modal fusion for human action recognition
D. He, F. Li, Q. Zhao, X. Long, Y. Fu, and S. Wen · 2018
Closest in time.
Human centric spatio-temporal action localization
J. Jiang, Y. Cao, L. Song, S. Z. Y. Li, Z. Xu, Q. Wu, C. Gan, C. Zhang, and G. Yu · 2018
Closest in time.
http://activity-net.org/challenges/2018/evaluation.html
Leaderboard:ActivityNet-AVA · 2018
Closest in time.
Actor-centric relation network
C. Sun, A. Shrivastava, C. Vondrick, K. Murphy, R. Sukthankar, and C. Schmid · 2018
Closest in time.
A closer look at spatiotemporal convolutions for action recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Closest in time.
Long-term temporal convolutions for action recognition
G. Varol, I. Laptev, and C. Schmid · 2018
Closest in time.
Appearance-and-relation networks for video classification
L. Wang, W. Li, W. Li, and L. Van Gool · 2018
Closest in time.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Closest in time.
Videos as space-time region graphs
X. Wang and A. Gupta · 2018
Closest in time.
Compressed video action recognition
C.-Y. Wu, M. Zaheer, H. Hu, R. Manmatha, A. J. Smola, and P. Krähenbühl · 2018
Closest in time.
Temporal relational reasoning in videos
B. Zhou, A. Andonian, A. Oliva, and A. Torralba · 2018
Closest in time.
ECO: efficient convolutional network for online video understanding
M. Zolfaghari, K. Singh, and T. Brox · 2018
Closest in time.
http://activity-net.org/challenges/2019/evaluation.html
ActivityNet-Challenge · 2019
Closest in time.
A short note on the kinetics-700 human action dataset
J. Carreira, E. Noland, C. Hillier, and A. Zisserman · 2019
Closest in time.
SlowFast networks for video recognition in ActivityNet challenge 2019
C. Feichtenhofer, H. Fan, J. Malik, and K. He · 2019
Closest in time.