Fetching the paper…
Reading the bibliography…
Generalizing over temporal variations is a prerequisite for effective action recognition in videos.
A short note on the Kinetics-700 human action dataset
Carreira, J., Noland, E., Hillier, C., Zisserman, A., 2019 · 1907
Earlier work this paper cites.
On a proposed analytical machine, in: The Origins of Digital Computers. Springer, pp. 73–87
Ludgate, P.E., 1982 · 1982
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., Schmidhuber, J., 1997 · 1997
Earlier work this paper cites.
X3D: Expanding architectures for efficient video recognition
Feichtenhofer, C., 2020 · 2004
Earlier work this paper cites.
Neighbourhood components analysis, in: Advances in neural information processing systems (NIPS), pp. 513–520
Goldberger, J., Hinton, G.E., Roweis, S.T., Salakhutdinov, R.R., 2005 · 2005
Earlier work this paper cites.
HMDB: A large video database for human motion recognition, in: International Conference on Computer Vision (ICCV), IEEE. pp. 2556–2563
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., Serre, T., 2011 · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A.R., Shah, M., 2012 · 2012
Earlier work this paper cites.
3D convolutional neural networks for human action recognition
Ji, S., Xu, W., Yang, M., Yu, K., 2013 · 2013
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos, in: Advances in Neural Information Processing Systems (NIPS), pp. 568–576
Simonyan, K., Zisserman, A., 2014 · 2014
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the Kinetics dataset, in: Computer Vision and Pattern Recognition (CVPR), IEEE. pp. 4724–4733
Carreira, J., Zisserman, A., 2017 · 2017
Cited alongside, same era.
Going deeper into action recognition: A survey
Herath, S., Harandi, M., Porikli, F., 2017 · 2017
Cited alongside, same era.
Learning spatio-temporal representation with pseudo-3D residual networks, in: International Conference on Computer Vision (ICCV), IEEE. pp. 5534–5542
Qiu, Z., Yao, T., Mei, T., 2017 · 2017
Cited alongside, same era.
Multi-fiber networks for video recognition, in: European Conference on Computer Vision (ECCV), pp. 352–367
Chen, Y., Kalantidis, Y., Li, J., Yan, S., Feng, J., 2018 · 2018
Cited alongside, same era.
Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet?, in: Computer Vision and Pattern Recognition (CVPR), IEEE. pp. 18–22
Hara, K., Kataoka, H., Satoh, Y., 2018 · 2018
Cited alongside, same era.
SlowFast networks for video recognition, in: International Conference on Computer Vision (ICCV), IEEE. pp. 6202–6211
Feichtenhofer, C., Fan, H., Malik, J., He, K., 2019 · 2019
Later among the works it cites.
TSM: Temporal shift module for efficient video understanding, in: International Conference on Computer Vision (ICCV), IEEE. pp. 7083–7093
Lin, J., Gan, C., Han, S., 2019 · 2019
Later among the works it cites.
Moments in time dataset: One million videos for event understanding
Monfort, M., Andonian, A., Zhou, B., Ramakrishnan, K., Bargal, S.A., Yan, T., Brown, L., Fan, Q., Gutfreund, D., Vondrick, C., et al., 2019 · 2019
Later among the works it cites.
Learning spatio-temporal representation with local and global diffusion, in: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE. pp. 12056–12065
Qiu, Z., Yao, T., Ngo, C.W., Tian, X., Mei, T., 2019 · 2019
Later among the works it cites.
Class feature pyramids for video explanation, in: International Conference on Computer Vision Workshop (ICCVW), IEEE. pp. 4255–4264
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention clusters: Purely attention based local feature integration for video classification, in: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE. pp. 7834–7843
Long, X., Gan, C., De Melo, G., Wu, J., Liu, X., Wen, S., 2018 · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition, in: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE. pp. 6450–6459
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., Paluri, M., 2018 · 2018
Cited alongside, same era.
Psanet: Point-wise spatial attention network for scene parsing, in: European Conference on Computer Vision (ECCV), pp. 267–283
Zhao, H., Zhang, Y., Liu, S., Shi, J., Change Loy, C., Lin, D., Jia, J., 2018 · 2018
Cited alongside, same era.
Temporal cycle-consistency learning, in: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE. pp. 1801–1810
Dwibedi, D., Aytar, Y., Tompson, J., Sermanet, P., Zisserman, A., 2019 · 2019
Cited alongside, same era.
Gather-excite: Exploiting feature context in convolutional neural networks, in: Advances in Neural Information Processing Systems (NIPS), pp. 9401–9411
Hu, J., Shen, L., Albanie, S., Sun, G., Vedaldi, A., 2018a
Cited in the paper.
Squeeze-and-excitation networks, in: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE. pp. 7132–7141
Hu, J., Shen, L., Sun, G., 2018b
Cited in the paper.
Analyzing human-human interactions: A survey
Stergiou, A., Poppe, R., 2019a
Cited in the paper.
Stergiou, A., Kapidis, G., Kalliatakis, G., Chrysoulas, C., Poppe, R., Veltkamp, R., 2019 · 2019
Later among the works it cites.
Video classification with channel-separated convolutional networks, in: International Conference on Computer Vision (ICCV), IEEE. pp. 5552–5561
Tran, D., Wang, H., Torresani, L., Feiszli, M., 2019 · 2019
Later among the works it cites.
Learning correspondence from the cycle-consistency of time, in: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE. pp. 2566–2576
Wang, X., Jabri, A., Efros, A.A., 2019 · 2019
Later among the works it cites.
HACS: Human action clips and segments dataset for recognition and temporal localization, in: International Conference on Computer Vision (ICCV), IEEE. pp. 8668–8678
Zhao, H., Torralba, A., Torresani, L., Yan, Z., 2019 · 2019
Later among the works it cites.
A multigrid method for efficiently training video models, in: Conference on Computer Vision and Pattern Recognition (CVPR), IEEE. pp. 153–162
Wu, C.Y., Girshick, R., He, K., Feichtenhofer, C., Krähenbühl, P., 2020 · 2020
Closest in time.