Fetching the paper…
Reading the bibliography…
We propose a soft attention based model for the task of action recognition in videos.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
The dynamic representation of scenes
R. A. Rensink · 2000
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and F.-F. Li · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng · 2011
Earlier work this paper cites.
Theano: new features and speed improvements
F. Bastien, P. Lamblin, R. Pascanu, J. Bergstra, I. J. Goodfellow, A. Bergeron, N. Bouchard, D. Warde-Farley, and Y. Bengio · 2012
Earlier work this paper cites.
Hybrid speech recognition with deep bidirectional LSTM
A. Graves, N. Jaitly, and A.-r. Mohamed · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and F.-F. Li · 2014
Earlier work this paper cites.
Beyond gaussian pyramid: Multi-skip feature stacking for action recognition
Z.-Z. Lan, M. Lin, X. Li, A. G. Hauptmann, and B. Raj · 2014
Earlier work this paper cites.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, and K. Kavukcuoglu · 2014
Earlier work this paper cites.
Action recognition with stacked fisher vectors
X. Peng, C. Zou, Y. Qiao, and Q. Peng · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
DL-SFA: deeply-learned slow feature analysis for action recognition
L. Sun, K. Jia, T.-H. Chan, Y. Fang, G. Wang, and S. Yan · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. V. Le · 2014
Cited alongside, same era.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. J. Mooney, and K. Saenko · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Closest in time.
Beyond short snippets: Deep networks for video classification
J. Y.-H. Ng, M. J. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Closest in time.
Object detection networks on convolutional feature maps
S. Ren, K. He, R. B. Girshick, X. Zhang, and J. Sun · 2015
Closest in time.
Unsupervised learning of video representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Closest in time.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Modeling video evolution for action recognition
B. Fernando, E. Gavves, J. Oramas, A. Ghodrati, and T. Tuytelaars · 2015
Cited alongside, same era.
M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu · 2015
Cited alongside, same era.
What do 15,000 object categories tell us about classifying and localizing actions?
M. Jain, J. C. v. Gemert, and C. G. M. Snoek · 2015
Cited alongside, same era.
Learning wake-sleep recurrent attention models
J. Ba, R. Grosse, R. Salakhutdinov, and B. Frey
Cited in the paper.
Closest in time.
Deep image: Scaling up image recognition
R. Wu, S. Yan, Y. Shan, Q. Dang, and G. Sun · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Closest in time.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Closest in time.
Every moment counts: Dense detailed labeling of actions in complex videos
S. Yeung, O. Russakovsky, N. Jin, M. Andriluka, G. Mori, and F.-F. Li · 2015
Closest in time.
Exploiting image-trained CNN architectures for unconstrained video classification
S. Zha, F. Luisier, W. Andrews, N. Srivastava, and R. Salakhutdinov · 2015
Closest in time.