Fetching the paper…
Reading the bibliography…
This paper describes our solution for the video recognition task of the Google Cloud and YouTube-8M Video Understanding Challenge that ranked the 3rd place.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
All about vlad
R. Arandjelovic and A. Zisserman · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Devnet: A deep event network for multimedia event detection and evidence recounting
C. Gan, N. Wang, Y. Yang, D.-Y. Yeung, and A. G. Hauptmann · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
C3D: Generic features for video analysis
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Youtube-8m: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Later among the works it cites.
Webly-supervised video recognition by mutually voting for relevant web images and web video frames
C. Gan, C. Sun, L. Duan, and B. Gong · 2016
Later among the works it cites.
You lead, we exceed: Labor-free video concept learning by jointly exploiting web videos and images
C. Gan, T. Yao, K. Yang, Y. Yang, and T. Mei · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Deep recurrent models with fast-forward connections for neural machine translation
J. Zhou, Y. Cao, X. Wang, P. Li, and W. Xu · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A discriminative CNN video representation for event detection
Z. Xu, Y. Yang, and A. G. Hauptmann · 2015
Cited alongside, same era.
Deep speaker: an end-to-end neural speaker embedding system
C. Li, X. Ma, B. Jiang, X. Li, X. Zhang, X. Liu, Y. Cao, A. Kannan, and Z. Zhu · 2017
Closest in time.
TS-LSTM and temporal-inception: Exploiting spatiotemporal dynamics for activity recognition
C.-Y. Ma, M.-H. Chen, Z. Kira, and G. AlRegib · 2017
Closest in time.