Fetching the paper…
Reading the bibliography…
Understanding and interpreting human actions is a long-standing challenge and a critical indicator of perception in artificial intelligence.
Johansson, G.: Visual perception of biological motion and a model for its analysis. Perception & psychophysics 14
1973
Earlier work this paper cites.
Grice, H.P.: Logic and conversation. In: Speech acts, pp. 41–58. Brill (1975)
1975
Earlier work this paper cites.
Woodward, A.L.: Infants selectively encode the goal object of an actor’s reach. Cognition 69
1998
Earlier work this paper cites.
Land, M., Mennie, N., Rusted, J.: The roles of vision and eye movements in the control of activities of daily living. Perception 28
1999
Earlier work this paper cites.
Baldwin, D.A., Baird, J.A.: Discerning intentions in dynamic human action. Trends in Cognitive Sciences 5
2001
Earlier work this paper cites.
Baldwin, D.A., Baird, J.A., Saylor, M.M., Clark, M.A.: Infants parse dynamic action. Child development 72
2001
Earlier work this paper cites.
Rubinstein, J.S., Meyer, D.E., Evans, J.E.: Executive control of cognitive processes in task switching. Journal of experimental psychology: human perception and performance 27
2001
Earlier work this paper cites.
Gergely, G., Bekkering, H., Király, I.: Rational imitation in preverbal infants. Nature 415
2002
Earlier work this paper cites.
Monsell, S.: Task switching. Trends in cognitive sciences 7
2003
Earlier work this paper cites.
Csibra, G., Gergely, G.: ‘obsessed with goals’: Functions and mechanisms of teleological interpretation of actions in humans. Acta psychologica (2007)
2007
Earlier work this paper cites.
Lerner, A., Chrysanthou, Y., Lischinski, D.: Crowds by example. In: Proceedings of Computer Graphics Forum (2007)
2007
Earlier work this paper cites.
Turaga, P., Chellappa, R., Subrahmanian, V.S., Udrea, O.: Machine recognition of human activities: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 18
2008
Earlier work this paper cites.
Choi, W., Shahid, K., Savarese, S.: What are they doing?: Collective activity classification using spatio-temporal relationship among people. In: International Conference on Computer Vision Workshops (ICCV Workshops) (2009)
2009
Earlier work this paper cites.
Pellegrini, S., Ess, A., Schindler, K., Van Gool, L.: You’ll never walk alone: Modeling social behavior for multi-target tracking. In: Proceedings of International Conference on Computer Vision (ICCV) (2009)
2009
Earlier work this paper cites.
Li, W., Zhang, Z., Liu, Z.: Action recognition based on a bag of 3d points. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2010)
2010
Earlier work this paper cites.
Oh, S., Hoogs, A., Perera, A., Cuntoor, N., Chen, C.C., Lee, J.T., Mukherjee, S., Aggarwal, J., Lee, H., Davis, L., et al.: A large-scale benchmark dataset for event recognition in surveillance video. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2011)
2011
Earlier work this paper cites.
Ryoo, M.S.: Human activity prediction: Early recognition of ongoing activities from streaming videos. In: Proceedings of International Conference on Computer Vision (ICCV) (2011)
2011
Earlier work this paper cites.
Pirsiavash, H., Ramanan, D.: Detecting activities of daily living in first-person camera views. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2012)
2012
Earlier work this paper cites.
Rohrbach, M., Amin, S., Andriluka, M., Schiele, B.: A database for fine grained activity detection of cooking activities. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2012)
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
Rohrbach, M., Amin, S., Andriluka, M., Schiele, B.: A database for fine grained activity detection of cooking activities. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2012)
2012
Earlier work this paper cites.
Ionescu, C., Papava, D., Olaru, V., Sminchisescu, C.: Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 36
2013
Earlier work this paper cites.
Koppula, H.S., Gupta, R., Saxena, A.: Learning human activities and object affordances from rgb-d videos. International Journal of Robotics Research (IJRR) 32
2013
Earlier work this paper cites.
Stein, S., McKenna, S.J.: Combining embedded accelerometers with computer vision for recognizing food preparation activities. In: ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (2013)
2013
Earlier work this paper cites.
Vondrick, C., Patterson, D., Ramanan, D.: Efficiently scaling up crowdsourced video annotation. International Journal of Computer Vision (IJCV) 101
2013
Earlier work this paper cites.
Vondrick, C., Patterson, D., Ramanan, D.: Efficiently scaling up crowdsourced video annotation. International Journal of Computer Vision (IJCV) 101
2013
Cited alongside, same era.
Bojanowski, P., Lajugie, R., Bach, F., Laptev, I., Ponce, J., Schmid, C., Sivic, J.: Weakly supervised action labeling in videos under ordering constraints. In: Proceedings of European Conference on Computer Vision (ECCV) (2014)
2014
Cited alongside, same era.
Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., Fei-Fei, L.: Large-scale video classification with convolutional neural networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2014)
2014
Cited alongside, same era.
Kuehne, H., Arslan, A., Serre, T.: The language of actions: Recovering the syntax and semantics of goal-directed human activities. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2014)
2014
Cited alongside, same era.
2017
Later among the works it cites.
Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al.: The ”something something” video database for learning and evaluating visual common sense. In: Proceedings of International Conference on Computer Vision (ICCV) (2017)
2017
Later among the works it cites.
He, K., Gkioxari, G., Dollár, P., Girshick, R.: Mask r-cnn. In: Proceedings of International Conference on Computer Vision (ICCV) (2017)
2017
Later among the works it cites.
Toyer, S., Cherian, A., Han, T., Gould, S.: Human pose forecasting via deep markov models. In: International Conference on Digital Image Computing: Techniques and Applications (DICTA) (2017)
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kuehne, H., Arslan, A., Serre, T.: The language of actions: Recovering the syntax and semantics of goal-directed human activities. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2014)
2014
Cited alongside, same era.
Pennington, J., Socher, R., Manning, C.D.: Glove: Global vectors for word representation. In: Proceedings of the conference on Empirical Methods in Natural Language Processing (EMNLP) (2014)
2014
Cited alongside, same era.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., Parikh, D.: Vqa: Visual question answering. In: Proceedings of International Conference on Computer Vision (ICCV) (2015)
2015
Cited alongside, same era.
Caba Heilbron, F., Escorcia, V., Ghanem, B., Carlos Niebles, J.: Activitynet: A large-scale video benchmark for human activity understanding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015)
2015
Cited alongside, same era.
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. In: Proceedings of Advances in Neural Information Processing Systems (NeurIPS) (2015)
2015
Cited alongside, same era.
Shu, T., Xie, D., Rothrock, B., Todorovic, S., Chun Zhu, S.: Joint inference of groups, events and human roles in aerial videos. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015)
2015
Cited alongside, same era.
Wu, C., Zhang, J., Savarese, S., Saxena, A.: Watch-n-patch: Unsupervised understanding of actions and relations. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015)
2015
Cited alongside, same era.
Wu, C., Zhang, J., Savarese, S., Saxena, A.: Watch-n-patch: Unsupervised understanding of actions and relations. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015)
2015
Cited alongside, same era.
Damen, D., Doughty, H., Maria Farinella, G., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al.: Scaling egocentric vision: The epic-kitchens dataset. In: Proceedings of European Conference on Computer Vision (ECCV) (2018)
2018
Later among the works it cites.
Fouhey, D.F., Kuo, W.c., Efros, A.A., Malik, J.: From lifestyle vlogs to everyday interactions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Later among the works it cites.
Gu, C., Sun, C., Ross, D.A., Vondrick, C., Pantofaru, C., Li, Y., Vijayanarasimhan, S., Toderici, G., Ricco, S., Sukthankar, R., et al.: Ava: A video dataset of spatio-temporally localized atomic visual actions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Later among the works it cites.
Huang, S., Qi, S., Zhu, Y., Xiao, Y., Xu, Y., Zhu, S.C.: Holistic 3d scene parsing and reconstruction from a single rgb image. In: Proceedings of European Conference on Computer Vision (ECCV) (2018)
2018
Later among the works it cites.
Li, Y., Liu, M., Rehg, J.M.: In the eye of beholder: Joint learning of gaze and actions in first person video. In: Proceedings of European Conference on Computer Vision (ECCV) (2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
Yeung, S., Russakovsky, O., Jin, N., Andriluka, M., Mori, G., Fei-Fei, L.: Every moment counts: Dense detailed labeling of actions in complex videos. International Journal of Computer Vision (IJCV) 126
2018
Later among the works it cites.
Zhou, L., Xu, C., Corso, J.J.: Towards automatic learning of procedures from web instructional videos. In: Proceedings of AAAI Conference on Artificial Intelligence (AAAI) (2018)
2018
Later among the works it cites.
Damen, D., Doughty, H., Maria Farinella, G., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al.: Scaling egocentric vision: The epic-kitchens dataset. In: Proceedings of European Conference on Computer Vision (ECCV) (2018)
2018
Later among the works it cites.
Gupta, A., Johnson, J., Fei-Fei, L., Savarese, S., Alahi, A.: Social gan: Socially acceptable trajectories with generative adversarial networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Later among the works it cites.
2019
Later among the works it cites.
Chen, Y., Huang, S., Yuan, T., Qi, S., Zhu, Y., Zhu, S.C.: Holistic++ scene understanding: Single-view 3d holistic scene parsing and human pose estimation with human-object interaction and physical commonsense. In: Proceedings of International Conference on Computer Vision (ICCV) (2019)
2019
Later among the works it cites.
Feichtenhofer, C., Fan, H., Malik, J., He, K.: Slowfast networks for video recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Later among the works it cites.
Monfort, M., Andonian, A., Zhou, B., Ramakrishnan, K., Bargal, S.A., Yan, T., Brown, L., Fan, Q., Gutfruend, D., Vondrick, C., et al.: Moments in time dataset: one million videos for event understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) (2019)
2019
Later among the works it cites.
Tang, Y., Ding, D., Rao, Y., Zheng, Y., Zhang, D., Zhao, L., Lu, J., Zhou, J.: Coin: A large-scale dataset for comprehensive instructional video analysis. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Later among the works it cites.
Vinyals, O., Babuschkin, I., Czarnecki, W.M., Mathieu, M., Dudzik, A., Chung, J., Choi, D.H., Powell, R., Ewalds, T., Georgiev, P., et al.: Grandmaster level in starcraft ii using multi-agent reinforcement learning. Nature 575
2019
Later among the works it cites.
Wu, C.Y., Feichtenhofer, C., Fan, H., He, K., Krahenbuhl, P., Girshick, R.: Long-term feature banks for detailed video understanding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Later among the works it cites.
Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., Mordatch, I.: Emergent tool use from multi-agent autocurricula. In: International Conference on Learning Representations (ICLR) (2020)
2020
Closest in time.
Girdhar, R., Ramanan, D.: Cater: A diagnostic dataset for compositional actions and temporal reasoning. International Conference on Learning Representations (ICLR) (2020)
2020
Closest in time.
Qi, S., Jia, B., Huang, S., Wei, P., Zhu, S.C.: A generalized earley parser for human activity parsing and prediction. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) (2020)
2020
Closest in time.