Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J., Zisserman, A.: · 2017
Later among the works it cites.
Rethinking spatiotemporal feature learning for video understanding
Original
Xie, S., Sun, C., Huang, J., Tu, Z., Murphy, K.: · 2017
Later among the works it cites.
What actions are needed for understanding human actions in videos?
Sigurdsson, G.A., Russakovsky, O., Gupta, A.: · 2017
Later among the works it cites.
Semi-supervised classification with graph convolutional networks
Kipf, T.N., Welling, M.: · 2017
Later among the works it cites.
The ”something something” video database for learning and evaluating visual common sense
Original
Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fründ, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., Memisevic, R.: · 2017
Later among the works it cites.
Temporal relational reasoning in videos
Zhou, B., Andonian, A., Torralba, A.: · 2017
Later among the works it cites.
Revisiting the effectiveness of off-the-shelf temporal modeling approaches for large-scale video classification
Original
Bian, Y., Gan, C., Liu, X., Li, F., Long, X., Li, Y., Qi, H., Zhou, J., Wen, S., Lin, Y.: · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., He, K.: · 2017
Later among the works it cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Qiu, Z., Yao, T., Mei, T.: · 2017
Later among the works it cites.
A simple neural network module for relational reasoning
Santoro, A., Raposo, D., Barrett, D.G., Malinowski, M., Pascanu, R., Battaglia, P., Lillicrap, T.: · 2017
Later among the works it cites.
Visual interaction networks
Watters, N., Tacchetti, A., Weber, T., Pascanu, R., Battaglia, P., Zoran, D.: · 2017
Later among the works it cites.
Dense and low-rank Gaussian CRFs using deep embeddings
Chandra, S., Usunier, N., Kokkinos, I.: · 2017
Later among the works it cites.
Segmentation-aware convolutional networks using local attention masks
Harley, A., Derpanis, K., Kokkinos, I.: · 2017
Later among the works it cites.
Learning affinity via spatial propagation networks
Liu, S., De Mello, S., Gu, J., Zhong, G., Yang, M.H., Kautz, J.: · 2017
Later among the works it cites.
The more you know: Using knowledge graphs for image classification
Marino, K., Salakhutdinov, R., Gupta, A.: · 2017
Later among the works it cites.
Joint discovery of object states and manipulation actions
Alayrac, J.B., Sivic, J., Laptev, I., Lacoste-Julien, S.: · 2017
Later among the works it cites.
Scc: Semantic context cascade for efficient action detection
Heilbron, F.C., Barrios, W., Escorcia, V., Ghanem, B.: · 2017
Later among the works it cites.
Temporal dynamic graph lstm for action-driven video object detection
Yuan, Y., Liang, X., Wang, X., Yeung, D.Y., Gupta, A.: · 2017
Later among the works it cites.
Mask R-CNN
He, K., Gkioxari, G., Dollár, P., Girshick, R.: · 2017
Later among the works it cites.
Feature pyramid networks for object detection
Lin, T.Y., Dollár, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: · 2017
Later among the works it cites.
The kinetics human action video dataset
Original
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al.: · 2017
Later among the works it cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Original
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., He, K.: · 2017
Later among the works it cites.
Asynchronous temporal fields for action recognition
Sigurdsson, G.A., Divvala, S., Farhadi, A., Gupta, A.: · 2017
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., Paluri, M.: · 2018
Closest in time.
Relation networks for object detection
Hu, H., Gu, J., Zhang, Z., Dai, J., Wei, Y.: · 2018
Closest in time.
Detecting and recognizing human-object intaractions
Gkioxari, G., Girshick, R., Dollár, P., He, K.: · 2018
Closest in time.
Attend and interact: Higher-order object interactions for video understanding
Ma, C.Y., Kadav, A., Melvin, I., Kira, Z., AlRegib, G., Graf, H.P.: · 2018
Closest in time.
Non-local neural networks
Wang, X., Girshick, R., Gupta, A., He, K.: · 2018
Closest in time.
Spatial temporal graph convolutional networks for skeleton-based action recognition
Yan, S., Xiong, Y., Lin, D.: · 2018
Closest in time.
Detectron
Girshick, R., Radosavovic, I., Gkioxari, G., Dollár, P., He, K.: · 2018
Closest in time.