Fetching the paper…
Reading the bibliography…
State-of-the-art methods for video action recognition commonly use an ensemble of two networks: the spatial stream, which takes RGB frames as input, and the temporal stream, which takes optical flow as input.
A duality based approach for realtime tv-l 1 optical flow
C. Zach, T. Pock, and H. Bischof · 2007
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Earlier work this paper cites.
Real-time action recognition with enhanced motion vector cnns
B. Zhang, L. Wang, Z. Wang, Y. Qiao, and H. Wang · 2016
Earlier work this paper cites.
Y. Bian, C. Gan, X. Liu, F. Li, X. Long, Y. Li, H. Qi, J. Zhou, S. Wen, and Y. Lin · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
Flownet 2.0: Evolution of optical flow estimation with deep networks
E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox · 2017
Cited alongside, same era.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Cited alongside, same era.
Learning spatio-temporal representation with pseudo-3d residual networks
Z. Qiu, T. Yao, and T. Mei · 2017
Cited alongside, same era.
On the integration of optical flow and action recognition
L. Sevilla-Lara, Y. Liao, F. Guney, V. Jampani, A. Geiger, and M. J. Black · 2017
Cited alongside, same era.
Asynchronous temporal fields for action recognition
G. A. Sigurdsson, S. K. Divvala, A. Farhadi, and A. Gupta · 2017
Cited alongside, same era.
Ava: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, S. Vijayanarasimhan, C. Pantofaru, D. A. Ross, G. Toderici, Y. Li, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik · 2018
Closest in time.
What makes a video a video: Analyzing temporal information in video understanding models and datasets
D.-A. Huang, V. Ramanathan, D. Mahajan, L. Torresani, M. Paluri, L. Fei-Fei, and J. C. Niebles · 2018
Closest in time.
Motion feature network: Fixed motion filter for action recognition
S. Lee, M. Lee, S. Son, G. Park, and N. Kwak · 2018
Closest in time.
Graph distillation for action detection with privileged modalities
Z. Luo, J.-T. Hsieh, L. Jiang, J. C. Niebles, and L. Fei-Fei · 2018
Closest in time.
Actionflownet: Learning motion representation for action recognition
J. Y.-H. Ng, J. Choi, J. Neumann, and L. S. Davis · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Revisiting unreasonable effectiveness of data in deep learning era
C. Sun, A. Shrivastava, S. Singh, and A. Gupta · 2017
Cited alongside, same era.
Convnet architecture search for spatiotemporal feature learning
D. Tran, J. Ray, Z. Shou, S.-F. Chang, and M. Paluri · 2017
Cited alongside, same era.
Hidden two-stream convolutional networks for action recognition
Y. Zhu, Z. Lan, S. Newsam, and A. G. Hauptmann · 2017
Cited alongside, same era.
A short note about kinetics-600
J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman · 2018
Cited alongside, same era.
End-to-end learning of motion representation for video understanding
L. Fan, W. Huang, S. E. Chuang Gan, B. Gong, and J. Huang · 2018
Cited alongside, same era.
Born again neural networks
T. Furlanello, Z. Lipton, M. Tschannen, L. Itti, and A. Anandkumar · 2018
Cited alongside, same era.
Im2flow: Motion hallucination from static images for action recognition
R. Gao, B. Xiong, and K. Grauman · 2018
Cited alongside, same era.
A. Piergiovanni and M. S. Ryoo · 2018
Closest in time.
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume
D. Sun, X. Yang, M.-Y. Liu, and J. Kautz · 2018
Closest in time.
Optical flow guided feature: A fast and robust motion representation for video action recognition
S. Sun, Z. Kuang, L. Sheng, W. Ouyang, and W. Zhang · 2018
Closest in time.
A closer look at spatiotemporal convolutions for action recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Closest in time.
Appearance-and-relation networks for video classification
L. Wang, W. Li, W. Li, and L. Van Gool · 2018
Closest in time.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Closest in time.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy · 2018
Closest in time.