Learning spatio-temporal representation with pseudo-3d residual networks
Qiu, Z., Yao, T., and Mei, T. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Relation networks for object detection
Hu, H., Gu, J., Zhang, Z., Dai, J., and Wei, Y. (2018) · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., and Paluri, M. (2018) · 2018
Cited alongside, same era.
Non-local neural networks
Wang, X., Girshick, R., Gupta, A., and He, K. (2018) · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Xie, S., Sun, C., Huang, J., Tu, Z., and Murphy, K. (2018) · 2018
Cited alongside, same era.
Gcnet: Non-local networks meet squeeze-excitation networks and beyond
Cao, Y., Xu, J., Lin, S., Wei, F., and Hu, H. (2019) · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Feichtenhofer, C., Fan, H., Malik, J., and He, K. (2019) · 2019
Cited alongside, same era.
Local relation networks for image recognition
Hu, H., Zhang, Z., Xie, Z., and Lin, S. (2019) · 2019
Cited alongside, same era.
Tsm: Temporal shift module for efficient video understanding
Lin, J., Gan, C., and Han, S. (2019) · 2019
Cited alongside, same era.
Video classification with channel-separated convolutional networks
Tran, D., Wang, H., Torresani, L., and Feiszli, M. (2019) · 2019
Cited alongside, same era.
Unilmv2: Pseudo-masked language models for unified language model pre-training
Bao, H., Dong, L., Wei, F., Wang, W., Yang, N., Liu, X., Wang, Y., Gao, J., Piao, S., Zhou, M., et al. (2020) · 2020
Cited alongside, same era.