Fetching the paper…
Reading the bibliography…
Recognizing when people have false beliefs is crucial for understanding their actions.
Theory of mind for a humanoid robot
B. Scassellati · 2002
Earlier work this paper cites.
Utility data annotation with amazon mechanical turk
A. Sorokin and D. Forsyth · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Temporal causality for the analysis of visual events
K. Prabhakar, S. Oh, P. Wang, G. D. Abowd, and J. M. Rehg · 2010
Earlier work this paper cites.
Following gaze: gaze-following behavior as a window into social cognition
S. V. Shepherd · 2010
Earlier work this paper cites.
Bayesian theory of mind: Modeling joint belief-desire attribution
C. L. Baker, R. R. Saxe, and J. B. Tenenbaum · 2011
Earlier work this paper cites.
Action recognition by dense trajectories
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2011
Earlier work this paper cites.
Theano: new features and speed improvements
F. Bastien, P. Lamblin, R. Pascanu, J. Bergstra, I. Goodfellow, A. Bergeron, N. Bouchard, D. Warde-Farley, and Y. Bengio · 2012
Earlier work this paper cites.
Learning to recognize daily actions using gaze
A. Fathi, Y. Li, and J. M. Rehg · 2012
Earlier work this paper cites.
Activity forecasting
K. M. Kitani, B. D. Ziebart, J. A. Bagnell, and M. Hebert · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
H. Pirsiavash and D. Ramanan · 2012
Earlier work this paper cites.
Neil: Extracting visual knowledge from web data
X. Chen, A. Shrivastava, and A. Gupta · 2013
Earlier work this paper cites.
Learning Spatio-Temporal Structure from RGB-D Videos for Human Activity Detection and Anticipation
H. S. Koppula and A. Saxena · 2013
Cited alongside, same era.
Inferring Dark Matter and Dark Energy from Videos
D. Xie, S. Todorovic, and S.-C. Zhu · 2013
Cited alongside, same era.
Bringing semantics into focus using visual abstraction
C. L. Zitnick and D. Parikh · 2013
Cited alongside, same era.
Predicting object dynamics in scenes
D. F. Fouhey and C. L. Zitnick · 2014
Cited alongside, same era.
Unsupervised domain adaptation by backpropagation
Y. Ganin and V. Lempitsky · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Simultaneous deep transfer across domains and tasks
E. Tzeng, J. Hoffman, T. Darrell, and K. Saenko · 2015
Later among the works it cites.
Learning common sense through visual abstraction
R. Vedantam, X. Lin, T. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Later among the works it cites.
Anticipating the future by watching unlabeled video
C. Vondrick, H. Pirsiavash, and A. Torralba · 2015
Later among the works it cites.
Galileo: Perceiving physical object properties by integrating a physics engine with deep learning
J. Wu, I. Yildirim, J. J. Lim, B. Freeman, and J. Tenenbaum · 2015
Later among the works it cites.
Yin and Yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Seeing the arrow of time
L. C. Pickup, Z. Pan, D. Wei, Y. Shih, C. Zhang, A. Zisserman, B. Scholkopf, and W. T. Freeman · 2014
Cited alongside, same era.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles · 2015
Cited alongside, same era.
We Are Humor Beings: Understanding and Predicting Visual Humor
A. Chandrasekaran, A. Kalyan, S. Antol, M. Bansal, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
HICO: A benchmark for recognizing human-object interactions in images
Y.-W. Chao, Z. Wang, Y. He, J. Wang, and J. Deng · 2015
Cited alongside, same era.
Represent and Infer Human Theory of Mind for Human-Robot Interaction
Y. Zhao, S. Holtzen, T. Gao, and S.-C. Zhu · 2015
Later among the works it cites.
Social LSTM: Human Trajectory Prediction in Crowded Spaces
A. Alahi, K. Goel, V. Ramanathan, A. Robicquet, L. Fei-Fei, and S. Savarese · 2016
Closest in time.
Learning aligned cross-modal representations from weakly aligned data
L. Castrejon, Y. Aytar, C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Closest in time.
Learning Physical Intuition of Block Towers by Example
A. Lerer, S. Gross, and R. Fergus · 2016
Closest in time.
The Curious Robot: Learning Visual Representations via Physical Interactions
L. Pinto, D. Gandhi, Y. Han, Y.-L. Park, and A. Gupta · 2016
Closest in time.
Information gathering actions over human internal state
D. Sadigh, S. S. Sastry, S. A. Seshia, and A. Dragan · 2016
Closest in time.
Stating the Obvious: Extracting Visual Common Sense Knowledge
M. Yatskar, V. Ordonez, and A. Farhadi · 2016
Closest in time.