Fetching the paper…
Reading the bibliography…
This paper describes our solution for the video recognition task of ActivityNet Kinetics challenge that ranked the 1st place.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Devnet: A deep event network for multimedia event detection and evidence recounting
C. Gan, N. Wang, Y. Yang, D.-Y. Yeung, and A. G. Hauptmann · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Earlier work this paper cites.
C3D: Generic features for video analysis
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
You lead, we exceed: Labor-free video concept learning by jointly exploiting web videos and images
C. Gan, T. Yao, K. Yang, Y. Yang, and T. Mei · 2016
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. V. Gool · 2016
Cited alongside, same era.
Xception: Deep learning with depthwise separable convolutions
F. Chollet · 2017
Cited alongside, same era.
Depthwise separable convolutions for neural machine translation
L. Kaiser, A. N. Gomez, and F. Chollet · 2017
Closest in time.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Closest in time.
Temporal modeling approaches for large-scale youtube-8m video understanding
F. Li, C. Gan, X. Liu, Y. Bian, X. Long, Y. Li, Z. Li, J. Zhou, and S. Wen · 2017
Closest in time.
A Structured Self-attentive Sentence Embedding
Z. Lin, M. Feng, C. Nogueira dos Santos, M. Yu, B. Xiang, B. Zhou, and Y. Bengio · 2017
Closest in time.
Attention Is All You Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin · 2017
Cited alongside, same era.
Cnn architectures for large-scale audio classification
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. J. Weiss, and K. Wilson · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Closest in time.