Fetching the paper…
Reading the bibliography…
Events in natural videos typically arise from spatio-temporal interactions between actors and objects and involve multiple co-occurring activities and object classes.
Koller, D., Weber, J., Huang, T., Malik, J., Ogasawara, G., Rao, B., Russell, S.: Towards robust automatic traffic scene analysis in real-time. In: IEEE Conference on Computer Vision and Pattern Recognition (1994)
1994
Earlier work this paper cites.
Marszalek, M., Schmid, C.: Semantic Hierarchies for Visual Object Recognition. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 1–7 (2007). https://doi.org/10.1109/CVPR.2007.383272
2007
Earlier work this paper cites.
Oliva, A., Torralba, A.: The role of context in object recognition. Trends in Cognitive Sciences 11
2007
Earlier work this paper cites.
Marszalek, M., Schmid, C.: Constructing Category Hierarchies for Visual Recognition. In: European Conference on Computer Vision. pp. 479–491. Springer-Verlag, Berlin, Heidelberg (2008)
2008
Earlier work this paper cites.
Choi, M.J., Lim, J.J., Torralba, A., Willsky, A.S.: Exploiting hierarchical context on a large database of object categories. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 129–136 (2010). https://doi.org/10.1109/CVPR.2010.5540221
2010
Earlier work this paper cites.
Liu, J., Kuipers, B., Savarese, S.: Recognizing human actions by attributes. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 3337–3344 (2011). https://doi.org/10.1109/CVPR.2011.5995353
2011
Earlier work this paper cites.
Koppula, H.S., Gupta, R., Saxena, A.: Learning human activities and object affordances from rgb-d videos. International Journal Robotics Research 32
2013
Earlier work this paper cites.
Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J.: Distributed Representations of Words and Phrases and their Compositionality. In: Neural Information Processing Systems, pp. 3111–3119 (2013)
2013
Earlier work this paper cites.
Prest, A., Ferrari, V., Schmid, C.: Explicit modeling of human-object interactions in realistic videos. IEEE Transactions on Pattern Analysis and Machine Intelligence (2013)
2013
Earlier work this paper cites.
Zhu, Y., Nayak, N.M., Roy-Chowdhury, A.K.: Context-aware modeling and recognition of activities in video. In: IEEE Conference on Computer Vision and Pattern Recognition (2013)
2013
Earlier work this paper cites.
Assari, S.M., Zamir, A.R., Shah, M.: Video classification using semantic concept co-occurrences. In: IEEE Conference on Computer Vision and Pattern Recognition (2014)
2014
Earlier work this paper cites.
Deng, J., Ding, N., Jia, Y., Frome, A., Murphy, K., Bengio, S., Li, Y., Neven, H., Adam, H.: Large-Scale Object Classification Using Label Relation Graphs. In: European Conference on Computer Vision. pp. 48–64. Lecture Notes in Computer Science (2014)
2014
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: European Conference on Computer Vision. pp. 740–755. Springer International Publishing, Cham (2014)
2014
Earlier work this paper cites.
Simonyan, K., Zisserman, A.: Two-stream convolutional networks for action recognition in videos. In: Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N.D., Weinberger, K.Q. (eds.) Neural Information Processing Systems. pp. 568–576. Curran Associates, Inc. (2014)
2014
Earlier work this paper cites.
Chéron, G., Laptev, I., Schmid, C.: P-CNN: Pose-Based CNN Features for Action Recognition. In: IEEE International Conference on Computer Vision. pp. 3218–3226 (2015). https://doi.org/10.1109/ICCV.2015.368
2015
Earlier work this paper cites.
Gkioxari, G., Girshick, R., Malik, J.: Contextual Action Recognition with R*CNN. In: IEEE International Conference on Computer Vision. pp. 1080–1088 (2015). https://doi.org/10.1109/ICCV.2015.129
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In: IEEE International Conference on Computer Vision (2015)
2015
Earlier work this paper cites.
Ramanathan, V., Li, C., Deng, J., Han, W., Li, Z., Gu, K., Song, Y., Bengio, S., Rossenberg, C., Fei-Fei, L.: Learning semantic relationships for better action retrieval in images. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 1100–1109 (2015). https://doi.org/10.1109/CVPR.2015.7298713
2015
Earlier work this paper cites.
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: Learning spatiotemporal features with 3d convolutional networks. In: IEEE International Conference on Computer Vision (2015)
2015
Earlier work this paper cites.
Wang, X., Ji, Q.: Video event recognition with deep hierarchical context model. In: IEEE Conference on Computer Vision and Pattern Recognition (2015)
2015
Earlier work this paper cites.
Zhou, Y., Ni, B., and, Tian, Q.: Interaction part mining: A mid-level approach for fine-grained action recognition. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 3323–3331 (2015). https://doi.org/10.1109/CVPR.2015.7298953
2015
Earlier work this paper cites.
Deng, Z., Vahdat, A., Hu, H., Mori, G.: Structure Inference Machines: Recurrent Neural Networks for Analyzing Relations in Group Activity Recognition. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 4772–4781 (2016). https://doi.org/10.1109/CVPR.2016.516
2016
Earlier work this paper cites.
Jain, A., Zamir, A.R., Savarese, S., Saxena, A.: Structural-RNN: Deep Learning on Spatio-Temporal Graphs. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 5308–5317 (2016)
2016
Earlier work this paper cites.
Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: Hollywood in homes: Crowdsourcing data collection for activity understanding. In: European Conference on Computer Vision. pp. 510–526 (2016)
2016
Earlier work this paper cites.
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Val Gool, L.: Temporal segment networks: Towards good practices for deep action recognition. In: European Conference on Computer Vision (2016)
2016
Cited alongside, same era.
Yatskar, M., Zettlemoyer, L., Farhadi, A.: Situation recognition: Visual semantic role labeling for image understanding. In: IEEE Conference on Computer Vision and Pattern Recognition (2016)
2016
Cited alongside, same era.
Carreira, J., Zisserman, A.: Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 4724–4733 (2017). https://doi.org/10.1109/CVPR.2017.502
2017
Cited alongside, same era.
Dave, A., Russakovsky, O., Ramanan, D.: Predictive-Corrective Networks for Action Detection. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 2067–2076 (2017). https://doi.org/10.1109/CVPR.2017.223
2017
Cited alongside, same era.
Piergiovanni, A., Ryoo, M.S.: Learning Latent Super-Events to Detect Multiple Activities in Videos. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 5304–5313 (2018). https://doi.org/10.1109/CVPR.2018.00556
2018
Later among the works it cites.
Qi, S., Wang, W., Jia, B., Shen, J., Zhu, S.C.: Learning human-object interactions by graph parsing neural networks. In: European Conference on Computer Vision. pp. 401–417 (2018)
2018
Later among the works it cites.
Schlichtkrull, M., Kipf, T.N., Bloem, P., Van Den Berg, R., Titov, I., Welling, M.: Modeling relational data with graph convolutional networks. In: European Semantic Web Conference. pp. 593–607. Springer (2018)
2018
Later among the works it cites.
Sun, C., Shrivastava, A., Vondrick, C., Murphy, K., Sukthankar, R., Schmid, C.: Actor-centric relation network. In: European Conference on Computer Vision. pp. 318–334 (2018)
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gilmer, J., Schoenholz, S.S., Riley, P.F., Vinyals, O., Dahl, G.E.: Neural Message Passing for Quantum Chemistry. In: International Conference on Machine learning. pp. 1263–1272 (2017)
2017
Cited alongside, same era.
Jiang, Y.G., Wu, Z., Wang, J., Xue, X., Chang, S.F.: Exploiting feature and class relationships in video categorization with regularized deep neural networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 40
2017
Cited alongside, same era.
Kipf, T.N., Welling, M.: Semi-supervised classification with graph convolutional networks. In: International Conference on Learning Representations (2017)
2017
Cited alongside, same era.
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., Bernstein, M.S., Fei-Fei, L.: Visual genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision 123
2017
Cited alongside, same era.
Lea, C., Flynn, M.D., Vidal, R., Reiter, A., Hager, G.D.: Temporal Convolutional Networks for Action Segmentation and Detection. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 1003–1012 (2017). https://doi.org/10.1109/CVPR.2017.113
2017
Cited alongside, same era.
Li, R., Tapaswi, M., Liao, R., Jia, J., Urtasun, R., Fidler, S.: Situation recognition with graph neural networks. In: IEEE International Conference on Computer Vision (2017)
2017
Cited alongside, same era.
Mavroudi, E., Tao, L., Vidal, R.: Deep Moving Poselets for Video Based Action Recognition. In: IEEE Winter Applications of Computer Vision Conference. pp. 111–120 (2017). https://doi.org/10.1109/WACV.2017.20
2017
Cited alongside, same era.
Sigurdsson, G.A., Divvala, S., Farhadi, A., Gupta, A.: Asynchronous Temporal Fields for Action Recognition. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 5650–5659 (2017). https://doi.org/10.1109/CVPR.2017.599
2017
Cited alongside, same era.
Veličković, P., Cucurull, G., Casanova, A., Romero, A., Liò, P., Bengio, Y.: Graph Attention Networks. International Conference on Learning Representations (2018)
2018
Later among the works it cites.
Wang, X., Gupta, A.: Videos as space-time region graphs. In: European Conference on Computer Vision. pp. 413–431 (2018)
2018
Later among the works it cites.
Zhou, B., Andonian, A., Oliva, A., Torralba, A.: Temporal relational reasoning in videos. In: Computer Vision – ECCV 2018. pp. 831–846 (2018)
2018
Later among the works it cites.
Zhou, L., Zhou, Y., Corso, J.J., Socher, R., Xiong, C.: End-to-end dense video captioning with masked transformer. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 8739–8748 (2018)
2018
Later among the works it cites.
Zitnik, M., Agrawal, M., Leskovec, J.: Modeling polypharmacy side effects with graph convolutional networks. Bioinformatics p. 457–466 (2018)
2018
Later among the works it cites.
Bajaj, M., Wang, L., Sigal, L.: G3raphground: Graph-based language grounding. In: IEEE International Conference on Computer Vision. pp. 4281–4290 (2019)
2019
Closest in time.
Chen, Y., Rohrbach, M., Yan, Z., Shuicheng, Y., Feng, J., Kalantidis, Y.: Graph-based global reasoning networks. In: IEEE Conference on Computer Vision and Pattern Recognition (2019)
2019
Closest in time.
Girdhar, R., Carreira, J., Doersch, C., Zisserman, A.: Video action transformer network. In: IEEE Conference on Computer Vision and Pattern Recognition (2019)
2019
Closest in time.
Gong, L., Cheng, Q.: Exploiting edge features for graph neural networks. In: IEEE Conference on Computer Vision and Pattern Recognition (2019)
2019
Closest in time.
Huang, H., Zhou, L., Zhang, W., Xu, C.: Dynamic Graph Modules for Modeling Higher-Order Interactions in Activity Recognition. In: British Machine Vision Conference (2019)
2019
Closest in time.
Junior, N.I.N., Hu, H., Zhou, G., Deng, Z., Liao, Z., Mori, G.: Structured Label Inference for Visual Understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence pp. 1–1 (2019). https://doi.org/10.1109/TPAMI.2019.2893215
2019
Closest in time.
Li, K., Zhang, Y., Li, K., Li, Y., Fu, Y.: Visual semantic reasoning for image-text matching. In: IEEE International Conference on Computer Vision (2019)
2019
Closest in time.
Nicolicioiu, A., Duta, I., Leordeanu, M.: Recurrent space-time graph neural networks. In: Neural Information Processing Systems (2019)
2019
Closest in time.
Piergiovanni, A.J., Ryoo, M.S.: Temporal gaussian mixture layer for videos. In: International Conference on Machine learning (2019)
2019
Closest in time.
Xiong, Y., Huang, Q., Guo, L., Zhou, H., Zhou, B., Lin, D.: A graph-based framework to bridge movies and synopses. In: IEEE International Conference on Computer Vision (2019)
2019
Closest in time.
Yu, W., Zhou, J., Yu, W., Liang, X., Xiao, N.: Heterogeneous graph learning for visual commonsense reasoning. In: Advances in Neural Information Processing Systems 32. pp. 2769–2779. Curran Associates, Inc. (2019)
2019
Closest in time.
Zhang, Y., Tokmakov, P., Hebert, M., Schmid, C.: A structured model for action detection. In: IEEE Conference on Computer Vision and Pattern Recognition (2019)
2019
Closest in time.
Zhou, L., Kalantidis, Y., Chen, X., Corso, J.J., Rohrbach, M.: Grounded video description. In: IEEE Conference on Computer Vision and Pattern Recognition (2019)
2019
Closest in time.
Ghosh, P., Yao, Y., Davis, L., Divakaran, A.: Stacked spatio-temporal graph convolutional networks for action segmentation. In: IEEE Winter Applications of Computer Vision Conference (2020)
2020
Closest in time.