Fetching the paper…
Reading the bibliography…
Events defined by the interaction of objects in a scene are often of critical importance; yet important events may have insufficient labeled examples to train a conventional deep model to generalize to future object appearance.
Exploiting human actions and object context for recognition tasks
D. J. Moore, I. A. Essa, and M. H. Hayes · 1999
Earlier work this paper cites.
What, where and who? classifying events by scene and object recognition
L.-J. Li and L. Fei-Fei · 2007
Earlier work this paper cites.
Hierarchical recognition of human activities interacting with objects
M. S. Ryoo and J. K. Aggarwal · 2007
Earlier work this paper cites.
A scalable approach to activity recognition based on object use
J. Wu, A. Osuntogun, T. Choudhury, M. Philipose, and J. M. Rehg · 2007
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
A. Gupta, A. Kembhavi, and L. S. Davis · 2009
Earlier work this paper cites.
Convolutional learning of spatio-temporal features
G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler · 2010
Earlier work this paper cites.
Learning spatiotemporal graphs of human activities
W. Brendel and S. Todorovic · 2011
Earlier work this paper cites.
Understanding egocentric activities
A. Fathi, A. Farhadi, and J. M. Rehg · 2011
Earlier work this paper cites.
Recognition using visual phrases
M. A. Sadeghi and A. Farhadi · 2011
Earlier work this paper cites.
A combined pose, object, and feature model for action understanding
B. Packer, K. Saenko, and D. Koller · 2012
Earlier work this paper cites.
Recognizing human-object interactions in still images by modeling the mutual context of objects and human poses
B. Yao and L. Fei-Fei · 2012
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Deepdriving: Learning affordance for direct perception in autonomous driving
C. Chen, A. Seff, A. Kornhauser, and J. Xiao · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
Finding action tubes
G. Gkioxari and J. Malik · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
S. Ren, K. He, R. B. Girshick, and J. Sun · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Cited alongside, same era.
End to end learning for self-driving cars
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, et al · 2016
Cited alongside, same era.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Cited alongside, same era.
Object level visual reasoning in videos
F. Baradel, N. Neverova, C. Wolf, J. Mille, and G. Mori · 2018
Closest in time.
Relational inductive biases, deep learning, and graph networks
P. W. Battaglia, J. B. Hamrick, V. Bapst, A. Sanchez-Gonzalez, V. Zambaldi, M. Malinowski, A. Tacchetti, D. Raposo, A. Santoro, R. Faulkner, et al · 2018
Closest in time.
Video action transformer network
R. Girdhar, J. Carreira, C. Doersch, and A. Zisserman · 2018
Closest in time.
Detectron
R. Girshick, I. Radosavovic, G. Gkioxari, P. Dollár, and K. He · 2018
Closest in time.
Mapping images to scene graphs with permutation-invariant structured prediction
R. Herzig, M. Raboh, G. Chechik, J. Berant, and A. Globerson · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semi-supervised classification with graph convolutional networks
T. N. Kipf and M. Welling · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Cited alongside, same era.
Towards weaklysupervised action localization
P. Weinzaepfel, X. Martin, and C. Schmid · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
The” something something” video database for learning and evaluating visual common sense
R. Goyal, S. E. Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Cited alongside, same era.
Mask R-CNN
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Cited alongside, same era.
Uncertainty-aware reinforcement learning for collision avoidance
G. Kahn, A. Villaflor, V. Pong, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
J. Johnson, A. Gupta, and L. Fei-Fei · 2018
Closest in time.
Referring relationships
R. Krishna, I. Chami, M. S. Bernstein, and L. Fei-Fei · 2018
Closest in time.
Graph networks as learnable physics engines for inference and control
A. Sanchez-Gonzalez, N. Heess, J. T. Springenberg, J. Merel, M. Riedmiller, R. Hadsell, and P. Battaglia · 2018
Closest in time.
Charades-ego: A large-scale dataset of paired third and first person videos
G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari · 2018
Closest in time.
Actor-centric relation network
C. Sun, A. Shrivastava, C. Vondrick, K. Murphy, R. Sukthankar, and C. Schmid · 2018
Closest in time.
Long-term temporal convolutions for action recognition
G. Varol, I. Laptev, and C. Schmid · 2018
Closest in time.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Closest in time.
Videos as space-time region graphs
X. Wang and A. Gupta · 2018
Closest in time.
Videos as space-time region graphs
X. Wang and A. Gupta · 2018
Closest in time.
Relational deep reinforcement learning
V. Zambaldi, D. Raposo, A. Santoro, V. Bapst, Y. Li, I. Babuschkin, K. Tuyls, D. Reichert, T. Lillicrap, E. Lockhart, et al · 2018
Closest in time.
Temporal relational reasoning in videos
B. Zhou, A. Andonian, A. Oliva, and A. Torralba · 2018
Closest in time.
Video Action Transformer Network
R. Girdhar, J. a. Carreira, C. Doersch, and A. Zisserman · 2019
Closest in time.
Crash to not crash: Learn to identify dangerous vehicles using a simulator
H. Kim, C. Suh, K. Lee, and G. Hwang · 2019
Closest in time.
Learning latent scene-graph representations for referring relationships
M. Raboh, R. Herzig, G. Chechik, J. Berant, and A. Globerson · 2019
Closest in time.