Fetching the paper…
Reading the bibliography…
Anticipating actions and objects before they start or appear is a difficult problem in computer vision with several real-world applications.
Leveraging archival video for building face datasets
D. Ramanan, S. Baker, and S. Kakade · 2007
Earlier work this paper cites.
80 million tiny images: A large data set for nonparametric object and scene recognition
A. Torralba et al · 2008
Earlier work this paper cites.
Learning actions from the web
N. Ikizler-Cinbis, R. G. Cinbis, and S. Sclaroff · 2009
Earlier work this paper cites.
Deep learning from temporal coherence in video
H. Mobahi, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
High five: Recognising human interactions in tv shows
A. Patron-Perez et al · 2010
Earlier work this paper cites.
A data-driven approach for event prediction
J. Yuen and A. Torralba · 2010
Earlier work this paper cites.
Multi-hypothesis motion planning for visual object tracking
H. Gong, J. Sim, M. Likhachev, and J. Shi · 2011
Earlier work this paper cites.
Track to the future: Spatio-temporal video segmentation with long-range motion cues
J. Lezama et al · 2011
Earlier work this paper cites.
Parsing video events with goal inference and intent prediction
M. Pei, Y. Jia, and S.-C. Zhu · 2011
Earlier work this paper cites.
Human activity prediction: Early recognition of ongoing activities from streaming videos
M. Ryoo · 2011
Earlier work this paper cites.
What makes paris look like paris?
C. Doersch, S. Singh, A. Gupta, J. Sivic, and A. A. Efros · 2012
Earlier work this paper cites.
Activity forecasting
K. M. Kitani, B. D. Ziebart, J. A. Bagnell, and M. Hebert · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky et al · 2012
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
H. Pirsiavash and D. Ramanan · 2012
Earlier work this paper cites.
Action bank: A high-level representation of activity in video
S. Sadanand and J. J. Corso · 2012
Earlier work this paper cites.
Watching unlabeled video helps learn new human actions from very few labeled snapshots
C.-Y. Chen and K. Grauman · 2013
Cited alongside, same era.
Neil: Extracting visual knowledge from web data
X. Chen, A. Shrivastava, and A. Gupta · 2013
Cited alongside, same era.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue et al · 2013
Cited alongside, same era.
Building high-level features using large scale unsupervised learning
Q. V. Le et al · 2013
Cited alongside, same era.
Inferring” dark matter” and” dark energy” from videos
D. Xie, S. Todorovic, and S.-C. Zhu · 2013
Cited alongside, same era.
Learning everything about anything: Webly-supervised visual concept learning
S. K. Divvala et al · 2014
Cited alongside, same era.
Seeing the arrow of time
L. C. Pickup et al · 2014
Later among the works it cites.
Video (language) modeling: a baseline for generative models of natural videos
M. Ranzato et al · 2014
Later among the works it cites.
Cnn features off-the-shelf: an astounding baseline for recognition
A. S. Razavian et al · 2014
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava et al · 2014
Later among the works it cites.
Predicting actions from static scenes
T.-H. Vu, C. Olsson, I. Laptev, A. Oliva, and J. Sivic · 2014
Later among the works it cites.
Patch to the future: Unsupervised visual prediction
J. Walker, A. Gupta, and M. Hebert · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Predicting object dynamics in scenes
D. F. Fouhey and C. L. Zitnick · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
G. E. Hinton, O. Vinyals, and J. Dean · 2014
Cited alongside, same era.
Max-margin early event detectors
M. Hoai and F. De la Torre · 2014
Cited alongside, same era.
Action-reaction: Forecasting the dynamics of human interaction
D.-A. Huang and K. M. Kitani · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, et al · 2014
Cited alongside, same era.
Reconstructing storyline graphs for image recommendation from web community photos
G. Kim and E. P. Xing · 2014
Cited alongside, same era.
Learning deep features for scene recognition using places database
B. Zhou et al · 2014
Later among the works it cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Closest in time.
THUMOS challenge: Action recognition with a large number of classes, 2015
A. Gorban et al · 2015
Closest in time.
Predicting the future behavior of a time-varying probability distribution
C. H. Lampert · 2015
Closest in time.
Unsupervised learning of video representations using lstm
N. Srivastava et al · 2015
Closest in time.
Dense optical flow prediction from a static image
J. Walker, A. Gupta, and M. Hebert · 2015
Closest in time.
Unsupervised learning of visual representations using videos
X. Wang and A. Gupta · 2015
Closest in time.
Exploiting image-trained cnn architectures for unconstrained video classification
S. Zha et al · 2015
Closest in time.