Fetching the paper…
Reading the bibliography…
A major stumbling block to progress in understanding basic human interactions, such as getting out of bed or opening a refrigerator, is lack of good training data.
The Hidden Dimension
E. T. Hall · 1966
Earlier work this paper cites.
Estimation of the size of a closed population when capture probabilities vary among animals
K. P. Burnham and W. S. Overton · 1978
Earlier work this paper cites.
The ecological approach to visual perception
J. Gibson · 1979
Earlier work this paper cites.
Distinctive Image Features from Scale-Invariant Keypoints
D. Lowe · 2004
Earlier work this paper cites.
Grouplet: A structured image representation for recognizing human and object interactions
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
What makes a chair a chair?
H. Grabner, J. Gall, and L. van Gool · 2011
Earlier work this paper cites.
From 3D scene geometry to human workspace
A. Gupta, S. Satkin, A. Efros, and M. Hebert · 2011
Earlier work this paper cites.
Hand detection using multiple proposals
A. Mittal, A. Zisserman, and P. H. S. Torr · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
A. Torralba and A. A. Efros · 2011
Earlier work this paper cites.
Scene semantics from long-term observation of people
V. Delaitre, D. Fouhey, I. Laptev, J. Sivic, A. Efros, and A. Gupta · 2012
Earlier work this paper cites.
Learning to recognize daily actions using gaze
A. Fathi, Y. Li, and J. M. Rehg · 2012
Earlier work this paper cites.
People watching: Human actions as a cue for single-view geometry
D. F. Fouhey, V. Delaitre, A. Gupta, A. A. Efros, I. Laptev, and J. Sivic · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
H. Pirsiavash and D. Ramanan · 2012
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Unstructured human activity detection from rgbd images
J. Sung, C. Ponce, B. Selman, and A. Saxena · 2012
Cited alongside, same era.
Hallucinated humans as the hidden context for labeling 3D scenes
Y. Jiang and A. Saxena · 2013
Cited alongside, same era.
Learning human activities and object affordances from RGB-D videos
H. S. Koppula, R. Gupta, and A. Saxena · 2013
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Robot learning manipulation action plans by “watching” unconstrained videos from the world wide web
Learning action maps of large environments via first-person vision
N. Rhinehart and K. M. Kitani · 2016
Later among the works it cites.
A multi-scale CNN for affordance segmentation in RGB images
A. Roy and S. Todorovic · 2016
Later among the works it cites.
Much ado about time: Exhaustive annotation of temporal data
G. A. Sigurdsson, O. Russakovsky, A. Farhadi, I. Laptev, and A. Gupta · 2016
Later among the works it cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Later among the works it cites.
Krishnacam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks
K. K. Singh, K. Fatahalian, and A. A. Efros · 2016
Later among the works it cites.
An uncertain future: Forecasting from static images using variational autoencoders
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Yang, Y. Li, C. Fermüller, and Y. Aloimonos · 2014
Cited alongside, same era.
Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions
S. Bambach, S. Lee, D. Crandall, and C. Yu · 2015
Cited alongside, same era.
Hico: A benchmark for recognizing human-object interactions in images
Y.-W. Chao, Z. Wang, Y. He, J. Wang, and J. Deng · 2015
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles · 2015
Cited alongside, same era.
S. Gupta and J. Malik · 2015
Cited alongside, same era.
How do we use our hands? discovering a diverse set of common grasps
D.-A. Huang, W.-C. Ma, M. Ma, and K. M. Kitani · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
J. Walker, C. Doersch, A. Gupta, and M. Hebert · 2016
Later among the works it cites.
Actions ~ transformations
X. Wang, A. Farhadi, and A. Gupta · 2016
Later among the works it cites.
Physics 101: Learning physical object properties from unlabeled videos
J. Wu, J. J. Lim, H. Zhang, J. B. Tenenbaum, and W. T. Freeman · 2016
Later among the works it cites.
Learning deep features for discriminative localization
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2016
Later among the works it cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Closest in time.
The ”something something” video database for learning and evaluating visual common sense
R. Goyal, S. E. Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic · 2017
Closest in time.
AVA: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, S. Vijayanarasimhan, C. Pantofaru, D. A. Ross, G. Toderici, Y. Li, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik · 2017
Closest in time.
Densely connected convolutional networks
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger · 2017
Closest in time.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Closest in time.
Dilated residual networks
F. Yu, V. Koltun, and T. Funkhouser · 2017
Closest in time.
Places: A 10 million image database for scene recognition
B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba · 2017
Closest in time.