Fetching the paper…
Reading the bibliography…
Given a scene, what is going to move, and in what direction will it move? Such a question could be considered a non-semantic form of action prediction.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Recognizing human actions: A local svm approach
I. Laptev, B. Caputo, et al · 2004
Earlier work this paper cites.
A data-driven approach for event prediction
J. Yuen and A. Torralba · 2010
Earlier work this paper cites.
A database and evaluation methodology for optical flow
S. Baker, D. Scharstein, J. Lewis, S. Roth, M. J. Black, and R. Szeliski · 2011
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Sift flow: Dense correspondence across scenes and its applications
C. Liu, J. Yuen, and A. Torralba · 2011
Earlier work this paper cites.
Activity forecasting
K. Kitani, B. Ziebart, D. Bagnell, and M. Hebert · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Anticipating human activities using object affordances for reactive robotic response
H. S. Koppula and A. Saxena · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Cited alongside, same era.
DeepFlow: Large displacement optical flow with deep matching
P. Weinzaepfel, J. Revaud, Z. Harchaoui, and C. Schmid · 2013
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
Predicting object dynamics in scenes
D. Fouhey and C. L. Zitnick · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
Max-margin early event detectors
M. Hoai and F. De la Torre · 2014
A hierarchical representation for future action prediction
T. Lan, T.-C. Chen, and S. Savarese · 2014
Later among the works it cites.
Déjà vu: Motion prediction in static images
S. L. Pintea, J. C. van Gemert, and A. W. Smeulders · 2014
Later among the works it cites.
Video (language) modeling: a baseline for generative models of natural videos
M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra · 2014
Later among the works it cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Action-reaction: Forecasting the dynamics of human interaction
D.-A. Huang and K. M. Kitani · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Discriminatively trained dense surface normal estimation
B. Z. L’ubor Ladickỳ and M. Pollefeys
Cited in the paper.
Patch to the future: Unsupervised visual prediction
J. Walker, A. Gupta, and M. Hebert · 2014
Later among the works it cites.
Designing deep networks for surface normal estimation
X. Wang, D. F. Fouhey, and A. Gupta · 2014
Later among the works it cites.
Panda: Pose aligned networks for deep attribute modeling
N. Zhang, M. Paluri, M. Ranzato, T. Darrell, and L. Bourdev · 2014
Later among the works it cites.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Closest in time.