Fetching the paper…
Reading the bibliography…
Feature shifts have been shown to be useful for action recognition with CNN-based models since Temporal Shift Module (TSM) was proposed.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: · 2009
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A.R., Shah, M.: · 2012
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., Fei-Fei, L.: · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K., Zisserman, A.: · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Donahue, J., Anne Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., Darrell, T.: · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., Sun, J.: · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J., Zisserman, A.: · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., Zisserman, A.: · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Qiu, Z., Yao, T., Mei, T.: · 2017
Earlier work this paper cites.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Hara, K., Kataoka, H., Satoh, Y.: · 2018
Earlier work this paper cites.
Shift: A zero flop, zero parameter alternative to spatial convolutions
Wu, B., Wan, A., Yue, X., Jin, P., Zhao, S., Golmant, N., Gholaminejad, A., Gonzalez, J., Keutzer, K.: · 2018
Cited alongside, same era.
S3d: Single shot multi-span detector via fully 3d convolutional network
Zhang, D., Dai, X., Wang, X., Wang, Y.F.: · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., Paluri, M.: · 2018
Cited alongside, same era.
Non-local neural networks
Wang, X., Girshick, R., Gupta, A., He, K.: · 2018
Cited alongside, same era.
Video action transformer network
Girdhar, R., Carreira, J., Doersch, C., Zisserman, A.: · 2019
Cited alongside, same era.
Tsm: Temporal shift module for efficient video understanding
Lin, J., Gan, C., Han, S.: · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., Sutskever, I.: · 2021
Later among the works it cites.
Vivit: A video vision transformer
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., Schmid, C.: · 2021
Later among the works it cites.
Vidtr: Video transformer without convolutions
Li, X., Zhang, Y., Liu, C., Shuai, B., Zhu, Y., Brattoli, B., Chen, H., Marsic, I., Tighe, J.: · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
All you need is a few shifts: Designing efficient convolutional neural networks for image classification
Chen, W., Xie, D., Zhang, Y., Pu, S.: · 2019
Cited alongside, same era.
Learnable gated temporal shift module for deep video inpainting”
Chang, Y.L., Liu, Z.Y., Lee, K.Y., Hsu, W.: · 2019
Cited alongside, same era.
Gate-Shift Networks for Video Action Recognition
Sudhakaran, S., Escalera, S., Lanz, O.: · 2020
Cited alongside, same era.
Rubiksnet: Learnable 3d-shift for efficient video action recognition
Fan, L., Buch, S., Wang, G., Cao, R., Zhu, Y., Niebles, J.C., Fei-Fei, L.: · 2020
Cited alongside, same era.
Later among the works it cites.
Is space-time attention all you need for video understanding?
Bertasius, G., Wang, H., Torresani, L.: · 2021
Later among the works it cites.
An image is worth 16x16 words, what is a video worth?
Sharir, G., Noy, A., Zelnik-Manor, L.: · 2021
Later among the works it cites.
In: Token Shift Transformer for Video Classification. Association for Computing Machinery, New York, NY, USA (2021) 917–925
Zhang, H., Hao, Y., Ngo, C.W · 2021
Later among the works it cites.
Tokenlearner: Adaptive space-time tokenization for videos
Ryoo, M.S., Piergiovanni, A., Arnab, A., Dehghani, M., Angelova, A.: · 2021
Later among the works it cites.
Space-time mixing attention for video transformer
Bulat, A., Perez-Rua, J.M., Sudhakaran, S., Martinez, B., Tzimiropoulos, G.: · 2021
Later among the works it cites.