Fetching the paper…
Reading the bibliography…
We present a convolution-free approach to video classification built exclusively on self-attention over space and time.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
Li, L., Chen, Y.-C., Cheng, Y., Gan, Z., Yu, L., and Liu, J · 2005
Earlier work this paper cites.
Visualizing data using t-SNE
van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L., Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Layer normalization
Ba, L. J., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
A short note about kinetics-600
Carreira, J., Noland, E., Banki-Horvath, A., Hillier, C., and Zisserman, A · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Relation networks for object detection
Hu, H., Gu, J., Zhang, Z., Dai, J., and Wei, Y · 2018
Earlier work this paper cites.
Resound: Towards action recognition without representation bias
Li, Y., Li, Y., and Vasconcelos, N · 2018
Earlier work this paper cites.
Scaling neural machine translation
Ott, M., Edunov, S., Grangier, D., and Auli, M · 2018
Earlier work this paper cites.
Image transformer
Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., Ku, A., and Tran, D · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., and Paluri, M · 2018
Earlier work this paper cites.
Non-local neural networks
Wang, X., Girshick, R. B., Gupta, A., and He, K · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Xie, S., Sun, C., Huang, J., Tu, Z., and Murphy, K · 2018
Earlier work this paper cites.
End-to-end dense video captioning with masked transformer
Zhou, L., Zhou, Y., Corso, J. J., Socher, R., and Xiong, C · 2018
Earlier work this paper cites.
Attention augmented convolutional networks
Bello, I., Zoph, B., Le, Q., Vaswani, A., and Shlens, J · 2019
Earlier work this paper cites.
Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution
Chen, Y., Fan, H., Xu, B., Yan, Z., Kalantidis, Y., Rohrbach, M., Yan, S., and Feng, J · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q., and Salakhutdinov, R · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
More is less: Learning efficient video representations by big-little network and depthwise temporal aggregation
Fan, Q., Chen, C.-F. R., Kuehne, H., Pistoia, M., and Cox, D · 2019
Cited alongside, same era.
Quantifying attention flow in transformers, 2020
Abnar, S. and Zuidema, W · 2020
Later among the works it cites.
Classifying, segmenting, and tracking object instances in video with mask propagation
Bertasius, G. and Torresani, L · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Later among the works it cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Later among the works it cites.
On the relationship between self-attention and convolutional layers
Cordonnier, J., Loukas, A., and Jaggi, M · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Feichtenhofer, C., Fan, H., Malik, J., and He, K · 2019
Cited alongside, same era.
Video action transformer network
Girdhar, R., Carreira, J., Doersch, C., and Zisserman, A · 2019
Cited alongside, same era.
Axial attention in multidimensional transformers
Ho, J., Kalchbrenner, N., Weissenborn, D., and Salimans, T · 2019
Cited alongside, same era.
Ccnet: Criss-cross attention for semantic segmentation
Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., and Liu, W · 2019
Cited alongside, same era.
Stm: Spatiotemporal and motion encoding for action recognition
Jiang, B., Wang, M., Gan, W., Wu, W., and Yan, J · 2019
Cited alongside, same era.
Multimodal transformer networks for end-to-end video-grounded dialogue systems
Le, H., Sahoo, D., Chen, N., and Hoi, S · 2019
Cited alongside, same era.
Tsm: Temporal shift module for efficient video understanding
Lin, J., Gan, C., and Han, S · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2020
Later among the works it cites.
Pyslowfast
Fan, H., Li, Y., Xiong, B., Lo, W.-Y., and Feichtenhofer, C · 2020
Later among the works it cites.
X3d: Expanding architectures for efficient video recognition
Feichtenhofer, C · 2020
Later among the works it cites.
Actor-transformers for group activity recognition
Gavrilyuk, K., Sanford, R., Javan, M., and Snoek, C. G. M · 2020
Later among the works it cites.
Motionsqueeze: Neural motion feature learning for video understanding
Kwon, H., Kim, M., Kwak, S., and Cho, M · 2020
Later among the works it cites.
D3d: Distilled 3d networks for video action recognition
Stroud, J., Ross, D., Sun, C., Deng, J., and Sukthankar, R · 2020
Later among the works it cites.
RAFT: recurrent all-pairs field transforms for optical flow
Teed, Z. and Deng, J · 2020
Later among the works it cites.
Axial-deeplab: Stand-alone axial-attention for panoptic segmentation
Wang, H., Zhu, Y., Green, B., Adam, H., Yuille, A. L., and Chen, L · 2020
Later among the works it cites.
Attentionnas: Spatiotemporal attention cell search for video classification
Wang, X., Xiong, X., Neumann, M., Piergiovanni, A. J., Ryoo, M. S., Angelova, A., Kitani, K. M., and Hua, W · 2020
Later among the works it cites.
Scaling autoregressive video models
Weissenborn, D., Täckström, O., and Uszkoreit, J · 2020
Later among the works it cites.
Bert representations for video question answering
Yang, Z., Garcia, N., Chu, C., Otani, M., Nakashima, Y., and Takemura, H · 2020
Later among the works it cites.
Exploring self-attention for image recognition
Zhao, H., Jia, J., and Koltun, V · 2020
Later among the works it cites.
Only time can tell: Discovering temporal data for temporal modeling
Sevilla-Lara, L., Zha, S., Yan, Z., Goswami, V., Feiszli, M., and Torresani, L · 2021
Closest in time.