Fetching the paper…
Reading the bibliography…
Understanding human actions is a key problem in computer vision.
A procedure for computing the k best solutions to discrete optimization problems and its application to the shortest path problem
E. L. Lawler · 1972
Earlier work this paper cites.
Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception
H. Wimmer and J. Perner · 1983
Earlier work this paper cites.
People thinking about thinking people: the role of the temporo-parietal junction in “theory of mind”
R. Saxe and N. Kanwisher · 2003
Earlier work this paper cites.
Action understanding as inverse planning
C. L. Baker, R. Saxe, and J. B. Tenenbaum · 2009
Earlier work this paper cites.
Cutting-plane training of structural svms
T. Joachims, T. Finley, and C.-N. J. Yu · 2009
Earlier work this paper cites.
Learning models for object recognition from natural language descriptions
J. Wang, K. Markert, and M. Everingham · 2009
Earlier work this paper cites.
A survey on vision-based human action recognition
R. Poppe · 2010
Earlier work this paper cites.
Kenlm: Faster and smaller language model queries
K. Heafield · 2011
Earlier work this paper cites.
Baby talk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
Diverse m-best solutions in markov random fields
D. Batra, P. Yadollahpour, A. Guzman-Rivera, and G. Shakhnarovich · 2012
Earlier work this paper cites.
Max-margin early event detectors
M. Hoai and F. De la Torre · 2012
Earlier work this paper cites.
Deep networks for predicting human intent with respect to objects
R. Kelley, L. Wigand, B. Hamilton, K. Browne, M. Nicolescu, and M. Nicolescu · 2012
Earlier work this paper cites.
Activity forecasting
K. M. Kitani, B. D. Ziebart, J. A. Bagnell, and M. Hebert · 2012
Earlier work this paper cites.
Human intent prediction using markov decision processes
C. McGhan, A. Nasir, and E. Atkins · 2012
Earlier work this paper cites.
Patch to the future: Unsupervised visual prediction
J. Walker, A. Gupta, and M. Hebert · 2012
Cited alongside, same era.
Neil: Extracting visual knowledge from web data
X. Chen, A. Shrivastava, and A. Gupta · 2013
Cited alongside, same era.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue et al · 2013
Cited alongside, same era.
Anticipating human activities using object affordances for reactive robotic response
H. S. Koppula and A. Saxena · 2013
Cited alongside, same era.
Exploiting language models to recognize unseen actions
D. Le, R. Bernardi, and J. Uijlings · 2013
Cited alongside, same era.
Language for learning complex human-object interactions
M. Patel, C. H. Ek, N. Kyriazis, A. Argyros, J. V. Miro, and D. Kragic · 2013
Cited alongside, same era.
Predicting object dynamics in scenes
D. F. Fouhey and C. L. Zitnick · 2014
Closest in time.
Visual persuasion: Inferring communicative intents of images
J. Joo, W. Li, F. F. Steen, , and S.-C. Zhu · 2014
Closest in time.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Closest in time.
Cnn features off-the-shelf: an astounding baseline for recognition
A. S. Razavian et al · 2014
Closest in time.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Closest in time.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dual coordinate solvers for large-scale structural svms
D. Ramanan · 2013
Cited alongside, same era.
Predicting human intention in visual observations of hand/object interactions
D. Song, N. Kyriazis, I. Oikonomidis, C. Papazov, A. Argyros, D. Burschka, and D. Kragic · 2013
Cited alongside, same era.
Inferring” dark matter” and” dark energy” from videos
D. Xie, S. Todorovic, and S.-C. Zhu · 2013
Cited alongside, same era.
Bringing semantics into focus using visual abstraction
C. L. Zitnick and D. Parikh · 2013
Cited alongside, same era.
Learning the visual interpretation of sentences
C. L. Zitnick, D. Parikh, and L. Vanderwende · 2013
Cited alongside, same era.
N-gram counts and language models from the common crawl
C. Buck, K. Heafield, and B. van Ooyen · 2014
Cited alongside, same era.
Closest in time.
Reasoning about object affordances in a knowledge base representation
Y. Zhu, A. Fathi, and L. Fei-Fei · 2014
Closest in time.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Closest in time.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
Skip-thought vectors
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler · 2015
Closest in time.
Where are they looking?
A. Recasens, A. Khosla, C. Vondrick, and A. Torralba · 2015
Closest in time.
Movieqa: Understanding stories in movies through question-answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.