Fetching the paper…
Reading the bibliography…
This paper presents the method that underlies our submission to the untrimmed video classification task of ActivityNet Challenge 2016.
Invited paper: Automatic speech recognition: History, methods and challenges
D. O’Shaughnessy · 2008
Earlier work this paper cites.
Image classification with the fisher vector: Theory and practice
J. Sánchez, F. Perronnin, T. Mensink, and J. J. Verbeek · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles · 2015
Earlier work this paper cites.
Devnet: A deep event network for multimedia event detection and evidence recounting
C. Gan, N. Wang, Y. Yang, D.-Y. Yeung, and A. G. Hauptmann · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Cited alongside, same era.
Fusing multi-stream deep networks for video classification
Z. Wu, Y. Jiang, X. Wang, H. Ye, X. Xue, and J. Wang · 2015
Cited alongside, same era.
Recognize complex events from static images by fusing deep channels
Y. Xiong, K. Zhu, D. Lin, and X. Tang · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Bag of visual words and fusion methods for action recognition: Comprehensive study and good practice
X. Peng, L. Wang, X. Wang, and Y. Qiao · 2016
Closest in time.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Closest in time.
Deep convolutional neural networks and data augmentation for acoustic event detection
N. Takahashi, M. Gygli, B. Pfister, and L. Van Gool · 2016
Closest in time.
Mofap: A multi-level representation for action recognition
L. Wang, Y. Qiao, and X. Tang · 2016
Closest in time.
Temporal segment networks: towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Real-time action recognition with enhanced motion vector cnns
B. Zhang, L. Wang, Z. Wang, Y. Qiao, and H. Wang
Cited in the paper.
Z. Zhu, J. H. Engel, and A. Hannun · 2016
Closest in time.