Fetching the paper…
Reading the bibliography…
Attentive video modeling is essential for action recognition in unconstrained videos due to their rich yet redundant information over space and time.
Rensink, R.A.: The dynamic representation of scenes. Visual Cognition 1
2000
Earlier work this paper cites.
Baccouche, M., Mamalet, F., Wolf, C., Garcia, C., Baskurt, A.: Sequential deep learning for human action recognition. In: International Workshop on Human Behavior Understanding (2011)
2011
Earlier work this paper cites.
Ji, S., Xu, W., Yang, M., Yu, K.: 3d convolutional neural networks for human action recognition. IEEE Transactions of Pattern Analysis and Machine Intelligence 35
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., Fei-Fei, L.: Large-scale video classification with convolutional neural networks. In: CVPR. pp. 1725–1732 (2014)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Simonyan, K., Zisserman, A.: Two-stream convolutional networks for action recognition in videos. In: Advances in neural information processing systems. pp. 568–576 (2014)
2014
Earlier work this paper cites.
Zeiler, M.D., Fergus, R.: Visualizing and understanding convolutional networks. In: European Conference on Computer Vision. pp. 818–833 (2014)
2014
Earlier work this paper cites.
Bahdanau, D., Cho, K., Bengio, Y.: Neural machine translation by jointly learning to align and translate. In: International Conference on Learning Representations (2015)
2015
Earlier work this paper cites.
Donahue, J., Anne Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., Darrell, T.: Long-term recurrent convolutional networks for visual recognition and description. In: IEEE Conference on Computer Vision and Pattern Recognition (2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Lee, C.Y., Xie, S., Gallagher, P., Zhang, Z., Tu, Z.: Deeply-supervised nets. In: Artificial Intelligence and Statistics. pp. 562–570 (2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: Learning spatiotemporal features with 3d convolutional networks. In: IEEE International Conference on Computer Vision (2015)
2015
Earlier work this paper cites.
Yue-Hei Ng, J., Hausknecht, M., Vijayanarasimhan, S., Vinyals, O., Monga, R., Toderici, G.: Beyond short snippets: Deep networks for video classification. In: CVPR. pp. 4694–4702 (2015)
2015
Earlier work this paper cites.
Sharma, S., Kiros, R., Salakhutdinov, R.: Action recognition using visual attention. In: International Conference on Learning Representations Workshop (2016)
2016
Earlier work this paper cites.
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Van Gool, L.: Temporal segment networks: Towards good practices for deep action recognition. In: European Conference on Computer Vision (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Cao, C., Zhang, Y., Wu, Y., Lu, H., Cheng, J.: Egocentric gesture recognition using recurrent 3d convolutional neural networks with spatiotemporal transformer modules. In: IEEE International Conference on Computer Vision. pp. 3763–3771 (2017)
2017
Earlier work this paper cites.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 6299–6308 (2017)
2017
Earlier work this paper cites.
Du, W., Wang, Y., Qiao, Y.: Recurrent spatial-temporal attention network for action recognition in videos. T-IP 27
2017
Earlier work this paper cites.
Gehring, J., Auli, M., Grangier, D., Yarats, D., Dauphin, Y.N.: Convolutional sequence to sequence learning. In: International Conference on Machine Learning (2017)
2017
Cited alongside, same era.
Girdhar, R., Ramanan, D.: Attentional pooling for action recognition. In: Advances on Neural Information Processing Systems. pp. 34–45 (2017)
2017
Cited alongside, same era.
Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al.: The “something something” video database for learning and evaluating visual common sense. In: IEEE International Conference on Computer Vision (2017)
2017
Cited alongside, same era.
Qiu, Z., Yao, T., Mei, T.: Learning spatio-temporal representation with pseudo-3D residual networks. In: IEEE International Conference on Computer Vision (2017)
2017
Cited alongside, same era.
Wang, X., Girshick, R., Gupta, A., He, K.: Non-local neural networks. In: IEEE Conference on Computer Vision and Pattern Recognition (2018)
2018
Later among the works it cites.
Wang, X., Gupta, A.: Videos as space-time region graphs. In: European Conference on Computer Vision (2018)
2018
Later among the works it cites.
Woo, S., Park, J., Lee, J.Y., So Kweon, I.: Cbam: Convolutional block attention module. In: European Conference on Computer Vision. pp. 3–19 (2018)
2018
Later among the works it cites.
Yang, Z., Li, Y., Yang, J., Luo, J.: Action recognition with spatio–temporal visual attention on skeleton image sequences. T-CSVT 29
2018
Later among the works it cites.
Zhang, Y., Cao, C., Cheng, J., Lu, H.: Egogesture: a new dataset and benchmark for egocentric hand gesture recognition. T-Multimedia 20
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Song, S., Lan, C., Xing, J., Zeng, W., Liu, J.: An end-to-end spatio-temporal attention model for human action recognition from skeleton data. In: AAAI (2017)
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Advances on Neural Information Processing Systems (2017)
2017
Cited alongside, same era.
Chang, X., Hospedales, T.M., Xiang, T.: Multi-level factorisation net for person re-identification. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 2109–2118 (2018)
2018
Cited alongside, same era.
Damen, D., Doughty, H., Maria Farinella, G., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al.: Scaling egocentric vision: The epic-kitchens dataset. In: European Conference on Computer Vision. pp. 720–736 (2018)
2018
Cited alongside, same era.
Furlanello, T., Lipton, Z.C., Tschannen, M., Itti, L., Anandkumar, A.: Born again neural networks. In: International Conference on Machine Learning (2018)
2018
Cited alongside, same era.
Gu, J., Hu, H., Wang, L., Wei, Y., Dai, J.: Learning region features for object detection. In: European Conference on Computer Vision (2018)
2018
Cited alongside, same era.
Hu, J., Shen, L., Sun, G.: Squeeze-and-excitation networks. In: IEEE International Conference on Computer Vision (2018)
2018
Cited alongside, same era.
Zhang, Y., Xiang, T., Hospedales, T.M., Lu, H.: Deep mutual learning. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 4320–4328 (2018)
2018
Later among the works it cites.
Zhou, B., Andonian, A., Oliva, A., Torralba, A.: Temporal relational reasoning in videos. In: European Conference on Computer Vision (2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: ACL. pp. 4171–4186 (2019)
2019
Later among the works it cites.
Fan, Q., Chen, C.F.R., Kuehne, H., Pistoia, M., Cox, D.: More is less: Learning efficient video representations by big-little network and depthwise temporal aggregation. In: Advances on Neural Information Processing Systems (2019)
2019
Later among the works it cites.
Feichtenhofer, C., Fan, H., Malik, J., He, K.: SlowFast networks for video recognition. In: IEEE International Conference on Computer Vision (2019)
2019
Later among the works it cites.
Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., Lu, H.: Dual attention network for scene segmentation. In: IEEE Conference on Computer Vision and Pattern Recognition (2019)
2019
Later among the works it cites.
Girdhar, R., Carreira, J., Doersch, C., Zisserman, A.: Video action transformer network. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 244–253 (2019)
2019
Later among the works it cites.
Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., Liu, W.: Ccnet: Criss-cross attention for semantic segmentation. In: IEEE International Conference on Computer Vision. pp. 603–612 (2019)
2019
Later among the works it cites.
Lin, J., Gan, C., Han, S.: Tsm: Temporal shift module for efficient video understanding. In: IEEE International Conference on Computer Vision. pp. 7083–7093 (2019)
2019
Later among the works it cites.
Martinez, B., Modolo, D., Xiong, Y., Tighe, J.: Action recognition with spatial-temporal discriminative filter banks. In: IEEE International Conference on Computer Vision (2019)
2019
Later among the works it cites.
Meng, L., Zhao, B., Chang, B., Huang, G., Sun, W., Tung, F., Sigal, L.: Interpretable spatio-temporal attention for video action recognition. In: IEEE International Conference on Computer Vision Workshop. pp. 0–0 (2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Si, C., Chen, W., Wang, W., Wang, L., Tan, T.: An attention enhanced graph convolutional lstm network for skeleton-based action recognition. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 1227–1236 (2019)
2019
Later among the works it cites.
Tran, D., Wang, H., Torresani, L., Feiszli, M.: Video classification with channel-separated convolutional networks. In: IEEE International Conference on Computer Vision (2019)
2019
Later among the works it cites.
Wu, C.Y., Feichtenhofer, C., Fan, H., He, K., Krahenbuhl, P., Girshick, R.: Long-term feature banks for detailed video understanding. In: IEEE Conference on Computer Vision and Pattern Recognition (2019)
2019
Later among the works it cites.