Fetching the paper…
Reading the bibliography…
Extracting temporal and representation features efficiently plays a pivotal role in understanding visual sequence information.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Recognizing human actions: a local svm approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Training and analysing deep recurrent neural networks
M. Hermans and B. Schrauwen · 2013
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
How to construct deep recurrent neural networks
R. Pascanu, C. Gulcehre, K. Cho, and Y. Bengio · 2013
Earlier work this paper cites.
Overfeat: Integrated recognition, localization and detection using convolutional networks
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Deep recursive neural networks for compositionality in language
O. Irsoy and C. Cardie · 2014
Earlier work this paper cites.
Opinion mining with deep recurrent neural networks
O. Irsoy and C. Cardie · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling
H. Sak, A. Senior, and F. Beaufays · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2014
Cited alongside, same era.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Encouraging lstms to anticipate actions very early
M. S. Aliakbarian, F. S. Saleh, M. Salzmann, B. Fernando, L. Petersson, and L. Andersson · 2017
Later among the works it cites.
Revisiting the effectiveness of off-the-shelf temporal modeling approaches for large-scale video classification
Y. Bian, C. Gan, X. Liu, F. Li, X. Long, Y. Li, H. Qi, J. Zhou, S. Wen, and Y. Lin · 2017
Later among the works it cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Later among the works it cites.
Annotating object instances with a polygon-rnn
L. Castrejon, K. Kundu, R. Urtasun, and S. Fidler · 2017
Later among the works it cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, S. Vijayanarasimhan, C. Pantofaru, D. A. Ross, G. Toderici, Y. Li, S. Ricco, R. Sukthankar, C. Schmid, et al · 2017
Later among the works it cites.
Mask r-cnn
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P.-S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Deep multi-scale video prediction beyond mean square error
M. Mathieu, C. Couprie, and Y. LeCun · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Learning to track for spatio-temporal action localization
P. Weinzaepfel, Z. Harchaoui, and C. Schmid · 2015
Cited alongside, same era.
Modeling spatial-temporal clues in a hybrid deep learning framework for video classification
Z. Wu, X. Wang, Y.-G. Jiang, H. Ye, and X. Xue · 2015
Cited alongside, same era.
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Later among the works it cites.
Tube convolutional neural network (t-cnn) for action detection in videos
R. Hou, C. Chen, and M. Shah · 2017
Later among the works it cites.
Decomposing motion and content for natural video sequence prediction
R. Villegas, J. Yang, S. Hong, X. Lin, and H. Lee · 2017
Later among the works it cites.
Efficient interactive annotation of segmentation datasets with polygon-rnn++
D. Acuna, H. Ling, A. Kar, and S. Fidler · 2018
Closest in time.
Crowdpose: Efficient crowded scenes pose estimation and a new benchmark
J. Li, C. Wang, H. Zhu, Y. Mao, H.-S. Fang, and C. Lu · 2018
Closest in time.
Transferable interactiveness prior for human-object interaction detection
Y.-L. Li, S. Zhou, X. Huang, L. Xu, Z. Ma, H.-S. Fang, Y.-F. Wang, and C. Lu · 2018
Closest in time.
Videolstm convolves, attends and flows for action recognition
Z. Li, K. Gavrilyuk, E. Gavves, M. Jain, and C. G. Snoek · 2018
Closest in time.
Attention clusters: Purely attention based local feature integration for video classification
X. Long, C. Gan, G. de Melo, J. Wu, X. Liu, and S. Wen · 2018
Closest in time.
Recurrent residual module for fast inference in videos
B. Pan, W. Lin, X. Fang, C. Huang, B. Zhou, and C. Lu · 2018
Closest in time.
Human action adverb recognition: Adha dataset and a three-stream hybrid model
B. Pang, K. Zha, and C. Lu · 2018
Closest in time.
Yolov3: An incremental improvement
J. Redmon and A. Farhadi · 2018
Closest in time.
Srda: Generating instance segmentation annotation via scanning, reasoning and domain adaptation
W. Xu, Y. Li, and C. Lu · 2018
Closest in time.