Fetching the paper…
Reading the bibliography…
Reducing redundancy is crucial for improving the efficiency of video recognition models.
L. Wang, Z. Tong, B. Ji, and G. Wu, “Tdn: Temporal difference networks for efficient action recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2021, pp. 1895–1904
1904
Earlier work this paper cites.
M. Bregonzio, S. Gong, and T. Xiang, “Recognising action as clouds of space-time interest points,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2009, pp. 1948–1955
1955
Earlier work this paper cites.
P. Dollár, V. Rabaud, G. Cottrell, and S. Belongie, “Behavior recognition via sparse spatio-temporal features,” in International Workshop on Visual Surveillance and Performance Evaluation of Tracking and Surveillance , 2005, pp. 65–72
2005
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2009, pp. 248–255
2009
Earlier work this paper cites.
B. Jiang, M. Wang, W. Gan, W. Wu, and J. Yan, “Stm: Spatiotemporal and motion encoding for action recognition,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2019, pp. 2000–2009
2009
Earlier work this paper cites.
H. Wang and C. Schmid, “Action recognition with improved trajectories,” in Proc. IEEE/CVF Comput. Vis. Pattern Recognit. , 2013, pp. 3551–3558
2013
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in Proc. Adv. Neural Inf. Process. Syst. , 2014, pp. 568–576
2014
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2015, pp. 4489–4497
2015
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in Proc. IEEE/CVF Comput. Vis. Pattern Recognit. , 2016, pp. 2921–2929
2016
Earlier work this paper cites.
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks: Towards good practices for deep action recognition,” in Proc. Eur. Conf. Comput. Vis. , 2016, pp. 20–36
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE/CVF Comput. Vis. Pattern Recognit. , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in Proc. IEEE/CVF Comput. Vis. Pattern Recognit. , 2017, pp. 6299–6308
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2017, pp. 618–626
2017
Earlier work this paper cites.
R. Goyal, S. E. Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag et al. , “The” something something” video database for learning and evaluating visual common sense.” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , vol. 1, no. 4, 2017, p. 5
2017
Earlier work this paper cites.
Z. Qiu, T. Yao, and T. Mei, “Learning spatio-temporal representation with pseudo-3d residual networks,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2017, pp. 5533–5541
2017
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in Proc. IEEE/CVF Comput. Vis. Pattern Recognit. , 2017, pp. 652–660
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh, “Realtime multi-person 2d pose estimation using part affinity fields,” in Proc. IEEE/CVF Comput. Vis. Pattern Recognit. , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst. , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
H. Fan, Z. Xu, L. Zhu, C. Yan, J. Ge, and Y. Yang, “Watching a small portion could be as good as watching all: Towards efficient video classification,” in Proc. Int. Joint Conf. Artif. Intell. , 2018
2018
Earlier work this paper cites.
Y. Li, Y. Li, and N. Vasconcelos, “Resound: Towards action recognition without representation bias,” in Proc. Eur. Conf. Comput. Vis. , 2018, pp. 513–528
2018
Earlier work this paper cites.
Y. Xu, Y. Han, R. Hong, and Q. Tian, “Sequential video vlad: Training the aggregation locally and temporally,” IEEE Trans. Image Process. , vol. 27, no. 10, pp. 4933–4944, 2018
2018
Earlier work this paper cites.
K. Hara, H. Kataoka, and Y. Satoh, “Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?” in Proc. IEEE/CVF Comput. Vis. Pattern Recognit. , 2018, pp. 6546–6555
2018
Earlier work this paper cites.
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,” in Proc. Eur. Conf. Comput. Vis. , 2018, pp. 305–321
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proc. IEEE/CVF Comput. Vis. Pattern Recognit. IEEE, 2018. [Online]. Available: http://dx.doi.org/10.1109/CVPR.2018.00745
2018
Cited alongside, same era.
X. Wang and A. Gupta, “Videos as space-time region graphs,” in Proc. Eur. Conf. Comput. Vis. , 2018, pp. 399–417
2018
Cited alongside, same era.
M. Zolfaghari, K. Singh, and T. Brox, “Eco: Efficient convolutional network for online video understanding,” in Proc. Eur. Conf. Comput. Vis. , 2018, pp. 695–712
S. Sudhakaran, S. Escalera, and O. Lanz, “Gate-shift networks for video action recognition,” in Proc. IEEE/CVF Comput. Vis. Pattern Recognit. , 2020, pp. 1102–1111
2020
Later among the works it cites.
Z. Wu, H. Li, C. Xiong, Y.-G. Jiang, and L. S. Davis, “A dynamic frame selection framework for fast video recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , 2020
2020
Later among the works it cites.
Y. Meng, C.-C. Lin, R. Panda, P. Sattigeri, L. Karlinsky, A. Oliva, K. Saenko, and R. Feris, “Ar-net: Adaptive frame resolution for efficient action recognition,” in Proc. Eur. Conf. COmput. Vis. Springer, 2020, pp. 86–104
2020
Later among the works it cites.
Z. Xie, Z. Zhang, X. Zhu, G. Huang, and S. Lin, “Spatially adaptive inference with stochastic feature sampling and interpolation,” in Proc. Eur. Conf. Comput. Vis. , August 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
B. Zhou, A. Andonian, A. Oliva, and A. Torralba, “Temporal relational reasoning in videos,” in Proc. Eur. Conf. Comput. Vis. , 2018, pp. 803–818
2018
Cited alongside, same era.
2018
Cited alongside, same era.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in Proc. IEEE Int. Conf. Comput. Vis. , 2019, pp. 6202–6211
2019
Cited alongside, same era.
J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” in Proc. IEEE Int. Conf. Comput. Vis. , 2019, pp. 7083–7093
2019
Cited alongside, same era.
W. Wu, D. He, X. Tan, S. Chen, and S. Wen, “Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2019, pp. 6222–6231
2019
Cited alongside, same era.
B. Korbar, D. Tran, and L. Torresani, “Scsampler: Sampling salient clips from video for efficient action recognition,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2019, pp. 6232–6242
2019
Cited alongside, same era.
Y. Ji, Y. Zhan, Y. Yang, X. Xu, F. Shen, and H. T. Shen, “A context knowledge map guided coarse-to-fine action recognition,” IEEE Trans. Image Process. , vol. 29, pp. 2742–2752, 2019
2019
Cited alongside, same era.
L. Shi, Y. Zhang, J. Cheng, and H. Lu, “Skeleton-based action recognition with multi-stream adaptive graph convolutional networks,” IEEE Trans. Image Process. , vol. 29, pp. 9532–9545, 2020
2020
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in Int. Conf. Learn. Represent. , 2020
2020
Later among the works it cites.
Z. Liu, D. Luo, Y. Wang, L. Wang, Y. Tai, C. Wang, J. Li, F. Huang, and T. Lu, “Teinet: Towards an efficient architecture for video recognition,” in Proc. AAAI Conf. Artif. Intell. , vol. 34, no. 07, 2020, pp. 11 669–11 676
2020
Later among the works it cites.
L. Fan, S. Buch, G. Wang, R. Cao, Y. Zhu, J. C. Niebles, and L. Fei-Fei, “Rubiksnet: Learnable 3d-shift for efficient video action recognition,” in Proc Eur. Conf. Comput. Vis. , 2020, pp. 505–521
2020
Later among the works it cites.
X. Li, Y. Wang, Z. Zhou, and Y. Qiao, “Smallbignet: Integrating core and contextual views for video classification,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2020, pp. 1092–1101
2020
Later among the works it cites.
Y. Wang, Z. Chen, H. Jiang, S. Song, Y. Han, and G. Huang, “Adaptive focus for efficient video recognition,” in Proc. IEEE Int. Conf. Comput. Vis. , October 2021
2021
Later among the works it cites.
X. Wang, L. Zhu, H. Wang, and Y. Yang, “Interactive prototype learning for egocentric action recognition,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2021, pp. 8168–8177
2021
Later among the works it cites.
X. Chen, C. Gao, C. Li, Y. Yang, and D. Meng, “Infrared action detection in the dark via cross-stream attention mechanism,” IEEE Trans. Multimedia , 2021
2021
Later among the works it cites.
X. Wang, L. Zhu, and Y. Yang, “T2vlad: global-local sequence alignment for text-video retrieval,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2021, pp. 5079–5088
2021
Later among the works it cites.
S. Kumawat, M. Verma, Y. Nakashima, and S. Raman, “Depthwise spatio-temporal stft convolutional neural networks for human action recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , 2021
2021
Later among the works it cites.
L. Zhu, H. Fan, Y. Luo, M. Xu, and Y. Yang, “Temporal cross-layer correlation mining for action recognition,” IEEE Trans. Multimedia , pp. 1–1, 2021
2021
Later among the works it cites.
A. Habibian, D. Abati, T. S. Cohen, and B. E. Bejnordi, “Skip-convolutions for efficient video processing,” in Proc. IEEE/CVF Comput. Vis. Pattern Recognit. , 2021, pp. 2695–2704
2021
Later among the works it cites.
Y. P. Yi Yang, Yueting Zhuang, “Multiple knowledge representation for big data ai: framework, application and case studies,” Frontiers of Information Technology & Electronic Engineering , vol. -1, no. -1, 2021
2021
Later among the works it cites.
C. Bian, W. Feng, L. Wan, and S. Wang, “Structural knowledge distillation for efficient skeleton-based action recognition,” IEEE Trans. Image Process. , vol. 30, pp. 2963–2976, 2021
2021
Later among the works it cites.
Z. Liu, L. Wang, W. Wu, C. Qian, and T. Lu, “Tam: Temporal adaptive module for video recognition,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2021, pp. 13 708–13 718
2021
Later among the works it cites.
X. Chen, C. Gao, F. Yang, X. Wang, Y. Yang, and Y. Han, “Video-to-image casting: A flatting method for video analysis,” in Proc. ACM Int. Conf. Multimedia , 2021, pp. 4958–4966
2021
Later among the works it cites.
Y. Meng, R. Panda, C.-C. Lin, P. Sattigeri, L. Karlinsky, K. Saenko, A. Oliva, and R. Feris, “Adafuse: Adaptive temporal fusion network for efficient action recognition,” in Proc. Int. Conf. Learn. Reprent. , 2021
2021
Later among the works it cites.
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in Proc. Int. Conf. Mach. Learn. , July 2021
2021
Later among the works it cites.
2021
Later among the works it cites.