Fetching the paper…
Reading the bibliography…
Transformer models have shown great success handling long-range interactions, making them a promising tool for modeling video.
M. McCloskey and N. J. Cohen, “Catastrophic interference in connectionist networks: The sequential learning problem,” in Psychology of learning and motivation , 1989
1989
Earlier work this paper cites.
R. P. Rao and D. H. Ballard, “Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects,” Nature Neuroscience , 1999
1999
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in CVPR , 2009
2009
Earlier work this paper cites.
Z. Zhang and D. Tao, “Slow feature analysis for human action recognition,” IEEE TPAMI , 2012
2012
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv , 2014
2014
Earlier work this paper cites.
X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-k. Wong, and W.-c. Woo, “Convolutional lstm network: A machine learning approach for precipitation nowcasting,” in NeurIPS , 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in NeurIPS , 2015
2015
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in ICCV , 2015
2015
Earlier work this paper cites.
X. Wang and A. Gupta, “Unsupervised learning of visual representations using videos,” in ICCV , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015
2015
Earlier work this paper cites.
X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-K. Wong, and W.-c. Woo, “Convolutional lstm network: A machine learning approach for precipitation nowcasting,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
A. Parikh, O. Täckström, D. Das, and J. Uszkoreit, “A decomposable attention model for natural language inference,” in Empirical Methods in Natural Language Processing , 2016
2016
Earlier work this paper cites.
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” in CVPR , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in CVPR , 2017
2017
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in CVPR , 2017
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in CVPR , 2017
2017
Earlier work this paper cites.
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotemporal feature learning for video understanding,” ECCV , 2017
2017
Earlier work this paper cites.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” NeurIPS , 2017
2017
Earlier work this paper cites.
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag et al. , “The “something something” video database for learning and evaluating visual common sense,” in ICCV , 2017
2017
Earlier work this paper cites.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in CVPR , 2017
2017
Earlier work this paper cites.
A. Gordo, J. Almazan, J. Revaud, and D. Larlus, “End-to-end learning of deep visual representations for image retrieval,” ICCV , 2017
2017
Earlier work this paper cites.
A. v. d. Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” arXiv , 2017
2017
Earlier work this paper cites.
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” TPAMI , 2017
2017
Earlier work this paper cites.
B. Mahasseni, M. Lam, and S. Todorovic, “Unsupervised video summarization with adversarial lstm networks,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2017, pp. 202–211
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV , 2017
2017
Earlier work this paper cites.
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi, “Inception-v4, inception-resnet and the impact of residual connections on learning,” in AAAI , 2017
2017
Earlier work this paper cites.
F. Mahdisoltani, G. Berger, W. Gharbieh, D. Fleet, and R. Memisevic, “On the effectiveness of task granularity for transfer learning,” arXiv , 2018
2018
Earlier work this paper cites.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR , 2018
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” in OpenAI Preprint , 2018
2018
Earlier work this paper cites.
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, “End-to-end dense video captioning with masked transformer,” in CVPR , 2018
2018
Earlier work this paper cites.
J. Fajtl, H. S. Sokeh, V. Argyriou, D. Monekosso, and P. Remagnino, “Summarizing videos with attention,” in ACCV , 2018
2018
Earlier work this paper cites.
P. Shaw, J. Uszkoreit, and A. Vaswani, “Self-attention with relative position representations,” in NAACL , 2018
2018
Earlier work this paper cites.
K.-M. Kim, S.-H. Choi, J.-H. Kim, and B.-T. Zhang, “Multimodal dual attention memory for video story question answering,” in ECCV , 2018
2018
Earlier work this paper cites.
F. Rodrigues and F. Pereira, “Deep learning from crowds,” in AAAI Conference on AI , 2018
2018
Earlier work this paper cites.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv , 2018
2018
Earlier work this paper cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in CVPR , 2018
2018
Earlier work this paper cites.
D.-A. Huang, V. Ramanathan, D. Mahajan, L. Torresani, M. Paluri, L. Fei-Fei, and J. C. Niebles, “What makes a video a video: Analyzing temporal information in video understanding models and datasets,” in CVPR , 2018
2018
Earlier work this paper cites.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in CVPR , 2018
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in CVPR , 2018
2018
Earlier work this paper cites.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in CVPR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Computational Linguistics , 2019
2019
Earlier work this paper cites.
R. Girdhar, J. Carreira, C. Doersch, and A. Zisserman, “Video action transformer network,” in CVPR , 2019
2019
Earlier work this paper cites.
D. Purwanto, R. Renanda Adhi Pramono, Y.-T. Chen, and W.-H. Fang, “Extreme low resolution action recognition with spatial-temporal multi-head self-attention and knowledge distillation,” in CVPR , 2019
2019
Earlier work this paper cites.
C. Sun, F. Baradel, K. Murphy, and C. Schmid, “Contrastive bidirectional transformer for temporal representation learning,” arXiv , 2019
2019
Earlier work this paper cites.
J. Carreira, E. Noland, C. Hillier, and A. Zisserman, “A short note on the kinetics-700 human action dataset,” arXiv , 2019
2019
Earlier work this paper cites.
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,” in ICCV , 2019
2019
Earlier work this paper cites.
Y.-T. Liu, Y.-J. Li, F.-E. Yang, S.-F. Chen, and Y.-C. F. Wang, “Learning hierarchical self-attention for video summarization,” in ICIP , 2019
2019
Earlier work this paper cites.
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in ICCV , 2019
2019
Earlier work this paper cites.
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-xl: Attentive language models beyond a fixed-length context,” in ACL , 2019
2019
Earlier work this paper cites.
H. Seong, J. Hyun, and E. Kim, “Video multitask transformer network,” in ICCV , 2019
2019
Earlier work this paper cites.
K. Fang, A. Toshev, L. Fei-Fei, and S. Savarese, “Scene memory transformer for embodied agents in long-horizon tasks,” in CVPR , 2019
2019
Earlier work this paper cites.
R. R. A. Pramono, Y.-T. Chen, and W.-H. Fang, “Hierarchical self-attention network for action localization in videos,” in CVPR , 2019
2019
Earlier work this paper cites.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in ICCV , 2019
2019
Earlier work this paper cites.
J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” in ICCV , 2019
2019
Earlier work this paper cites.
Q. Xu, M. Zhang, Z. Gu, and G. Pan, “Overfitting remedy by sparsifying regularization on fully-connected layers of cnns,” Neurocomputing , 2019
2019
Earlier work this paper cites.
P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens, “Stand-alone self-attention in vision models,” in NeurIPS , 2019
2019
Earlier work this paper cites.
X. Chen, D. Liu, C. Lei, R. Li, Z.-J. Zha, and Z. Xiong, “Bert4sessrec: Content-based video relevance prediction with bidirectional encoder representations from transformer,” in ACM-MM , 2019
2019
Earlier work this paper cites.
D. Purwanto, R. R. A. Pramono, Y.-T. Chen, and W.-H. Fang, “Three-stream network with bidirectional self-attention for action recognition in extreme low resolution videos,” IEEE Signal Processing Letters , 2019
2019
Earlier work this paper cites.
D. Hendrycks, M. Mazeika, S. Kadavath, and D. Song, “Using self-supervised learning can improve model robustness and uncertainty,” NeurIPS , 2019
2019
Earlier work this paper cites.
T. Han, W. Xie, and A. Zisserman, “Video representation learning by dense predictive coding,” in ICCV-W , 2019
2019
Earlier work this paper cites.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in ICCV , 2019
2019
Earlier work this paper cites.
D. Ghadiyaram, D. Tran, and D. Mahajan, “Large-scale weakly-supervised pre-training for video action recognition,” in CVPR , 2019
2019
Earlier work this paper cites.
D. Tran, H. Wang, L. Torresani, and M. Feiszli, “Video classification with channel-separated convolutional networks,” in CVPR , 2019
2019
Earlier work this paper cites.
S. Serrano and N. A. Smith, “Is attention interpretable?” in ACL , 2019
2019
Earlier work this paper cites.
S. Jain and B. C. Wallace, “Attention is not explanation,” in NAACL-HLT , 2019
2019
Earlier work this paper cites.
S. Wiegreffe and Y. Pinter, “Attention is not not explanation,” in EMNLP-IJCNLP , 2019
2019
Earlier work this paper cites.
D. Kim, D. Cho, and I. S. Kweon, “Self-supervised video representation learning with space-time cubic puzzles,” in AAAI , 2019
2019
Earlier work this paper cites.
G. Kordopatis-Zilos, S. Papadopoulos, I. Patras, and I. Kompatsiaris, “Visil: Fine-grained spatio-temporal video similarity learning,” in ICCV , 2019
2019
Earlier work this paper cites.
J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” in ICCV , 2019
2019
Earlier work this paper cites.
A. Howard, M. Sandler, G. Chu, and L.-C. e. a. Chen, “Searching for mobilenetv3,” in CVPR , 2019
2019
Earlier work this paper cites.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” in ICLR , 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” in OpenAI Blog , 2019
2019
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV , 2020
2020
Earlier work this paper cites.
L. Zhu and Y. Yang, “Actbert: Learning global-local video-text representations,” in CVPR , 2020
2020
Earlier work this paper cites.
L. Li, Y.-C. Chen, Y. Cheng, Z. Gan, L. Yu, and J. Liu, “Hero: Hierarchical encoder for video+ language omni-representation pre-training,” in EMNLP , 2020
2020
Earlier work this paper cites.
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient transformers: A survey,” ACM CSUR , 2020
2020
Earlier work this paper cites.
K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu, Z. Yang, Y. Zhang, and D. Tao, “A survey on vision transformer,” in IEEE TPAMI , 2020
2020
Cited alongside, same era.
S. Ging, M. Zolfaghari, H. Pirsiavash, and T. Brox, “Coot: Cooperative hierarchical transformer for video-text representation learning,” in NeurIPS , 2020
2020
Cited alongside, same era.
M. E. Kalfaoglu, S. Kalkan, and A. A. Alatan, “Late temporal modeling in 3d cnn architectures with bert for action recognition,” in ECCV , 2020
2020
Cited alongside, same era.
J. Lei, L. Wang, Y. Shen, D. Yu, T. L. Berg, and M. Bansal, “Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning,” in ACL , 2020
2020
Cited alongside, same era.
Y. Gu, L. Wang, Z. Wang, Y. Liu, M.-M. Cheng, and S.-P. Lu, “Pyramid constrained self-attention network for fast video salient object detection,” in AAAI , 2020
H. Zhou, A. Kadav, F. Lai, A. Niculescu-Mizil, M. R. Min, M. Kapadia, and H. P. Graf, “Hopper: Multi-hop transformer for spatiotemporal reasoning,” in ICLR , 2021
2021
Later among the works it cites.
M. Xu, Y. Xiong, H. Chen, X. Li, W. Xia, Z. Tu, and S. Soatto, “Long short-term transformer for online action detection,” NeurIPS , 2021
2021
Later among the works it cites.
A. Bozic, P. Palafox, J. Thies, A. Dai, and M. Nießner, “Transformerfusion: Monocular rgb scene reconstruction using transformers,” NeurIPS , 2021
2021
Later among the works it cites.
W. Yu, H. Zheng, M. Li, L. Ji, L. Wu, N. Xiao, and N. Duan, “Learning from inside: Self-driven siamese sampling and reasoning for video question answering,” NeurIPS , 2021
2021
Later among the works it cites.
J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu, “Less is more: Clipbert for video-and-language learning via sparse sampling,” in CVPR , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
S. Kondo, “Lapformer: surgical tool detection in laparoscopic surgical video using transformer architecture,” Computer Methods in Biomechanics and Biomedical Engineering: Imaging and Visualization , 2020
2020
Cited alongside, same era.
A. Johnston and G. Carneiro, “Self-supervised monocular trained depth estimation using self-attention and discrete disparity volume,” in CVPR , 2020
2020
Cited alongside, same era.
X. Wang, X. Xiong, M. Neumann, A. Piergiovanni, M. S. Ryoo, A. Angelova, K. M. Kitani, and W. Hua, “Attentionnas: Spatiotemporal attention cell search for video classification,” in ECCV , 2020
2020
Cited alongside, same era.
Y. Zeng, J. Fu, and H. Chao, “Learning joint spatial-temporal transformations for video inpainting,” in ECCV , 2020
2020
Cited alongside, same era.
K. Gavrilyuk, R. Sanford, M. Javan, and C. G. Snoek, “Actor-transformers for group activity recognition,” in CVPR , 2020
2020
Cited alongside, same era.
V. Gabeur, C. Sun, K. Alahari, and C. Schmid, “Multi-modal Transformer for Video Retrieval,” in ECCV , 2020
2020
Cited alongside, same era.
R. Rakhimov, D. Volkhonskiy, A. Artemov, D. Zorin, and E. Burnaev, “Latent video transformer,” arXiv , 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
L. Sevilla-Lara, S. Zha, Z. Yan, V. Goswami, M. Feiszli, and L. Torresani, “Only time can tell: Discovering temporal data for temporal modeling,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2021
2021
Later among the works it cites.
C.-F. R. Chen, R. Panda, K. Ramakrishnan, R. Feris, J. Cohn, A. Oliva, and Q. Fan, “Deep analysis of cnn-based spatio-temporal representations for action recognition,” in CVPR , 2021
2021
Later among the works it cites.
X. Chen, C.-J. Hsieh, and B. Gong, “When vision transformers outperform resnets without pretraining or strong data augmentations,” arXiv , 2021
2021
Later among the works it cites.
Z. Yuan, X. Song, L. Bai, Z. Wang, and W. Ouyang, “Temporal-channel transformer for 3d lidar-based video object detection for autonomous driving,” Circuits and Systems for Video Technology , 2021
2021
Later among the works it cites.
X. Chen, B. Yan, J. Zhu, D. Wang, X. Yang, and H. Lu, “Transformer tracking,” in CVPR , 2021
2021
Later among the works it cites.
M. Patrick, P.-Y. Huang, I. Misra, F. Metze, A. Vedaldi, Y. M. Asano, and J. a. F. Henriques, “Space-time crop & attend: Improving cross-modal video representation learning,” in ICCV , 2021
2021
Later among the works it cites.
Y. Wang, Z. Xu, X. Wang, C. Shen, B. Cheng, H. Shen, and H. Xia, “End-to-end video instance segmentation with transformers,” in CVPR , 2021
2021
Later among the works it cites.
Z. Dai, H. Liu, Q. V. Le, and M. Tan, “Coatnet: Marrying convolution and attention for all data sizes,” in NeurIPS , 2021
2021
Later among the works it cites.
C. Sun, A. Nagrani, Y. Tian, and C. Schmid, “Composable augmentation encoding for video representation learning,” in ICCV , 2021
2021
Later among the works it cites.
Y. Chen and J. Joo, “Understanding and mitigating annotation bias in facial expression recognition,” in ICCV , 2021
2021
Later among the works it cites.
D. Kim, Y. Yoo, S. Park, J. Kim, and J. Lee, “Selfreg: Self-supervised contrastive regularization for domain generalization,” in ICCV , 2021
2021
Later among the works it cites.
X. Chen and K. He, “Exploring simple siamese representation learning,” in CVPR , 2021
2021
Later among the works it cites.
R. Qian, T. Meng, B. Gong, M.-H. Yang, H. Wang, S. Belongie, and Y. Cui, “Spatiotemporal contrastive video representation learning,” in CVPR , 2021
2021
Later among the works it cites.
C. Feichtenhofer, H. Fan, B. Xiong, R. Girshick, and K. He, “A large-scale study on unsupervised spatiotemporal representation learning,” in CVPR , 2021
2021
Later among the works it cites.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in ICCV , 2021
2021
Later among the works it cites.
Y. Tian, X. Chen, and S. Ganguli, “Understanding self-supervised learning dynamics without contrastive pairs,” in ICML , 2021
2021
Later among the works it cites.
J. Shao, X. Wen, B. Zhao, and X. Xue, “Temporal context aggregation for video retrieval with contrastive learning,” in WACV , 2021
2021
Later among the works it cites.
S. Liu, H. Fan, S. Qian, Y. Chen, W. Ding, and Z. Wang, “Hit: Hierarchical transformer with momentum contrast for video-text retrieval,” in ICCV , 2021
2021
Later among the works it cites.
S. Li, X. Li, J. Lu, and J. Zhou, “Self-supervised video hashing via bidirectional transformers,” in CVPR , 2021
2021
Later among the works it cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in ICML , 2021
2021
Later among the works it cites.
B. Zhang, J. Yu, C. Fifty, W. Han, A. M. Dai, R. Pang, and F. Sha, “Co-training transformer with videos and images improves action recognition,” arXiv , 2021
2021
Later among the works it cites.
Y. Dong, J.-B. Cordonnier, and A. Loukas, “Attention is not all you need: Pure attention loses rank doubly exponentially with depth,” in ICML , 2021
2021
Later among the works it cites.
S. Bhojanapalli, A. Chakrabarti, D. Glasner, D. Li, T. Unterthiner, and A. Veit, “Understanding robustness of transformers for image classification,” in ICCV , 2021
2021
Later among the works it cites.
K. Mahmood, R. Mahmood, and M. Van Dijk, “On the robustness of vision transformers to adversarial examples,” in ICCV , 2021
2021
Later among the works it cites.
M. M. Naseer, K. Ranasinghe, S. H. Khan, M. Hayat, F. Shahbaz Khan, and M.-H. Yang, “Intriguing properties of vision transformers,” NeurIPS , 2021
2021
Later among the works it cites.
J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny, “Barlow twins: Self-supervised learning via redundancy reduction,” in ICML , 2021
2021
Later among the works it cites.
M. Dzabraev, M. Kalashnikov, S. Komkov, and A. Petiushko, “Mdmmt: Multidomain multimodal transformer for video retrieval,” in CVPR , 2021
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” arXiv , 2021
2021
Later among the works it cites.
B. Yu, M. Tang, L. Zheng, G. Zhu, J. Wang, H. Feng, X. Feng, and H. Lu, “High-performance discriminative tracking with transformers,” in ICCV , 2021
2021
Later among the works it cites.
B. Yan, H. Peng, J. Fu, D. Wang, and H. Lu, “Learning spatio-temporal transformer for visual tracking,” in ICCV , 2021
2021
Later among the works it cites.
Z. Yuan, X. Song, L. Bai, Z. Wang, and W. Ouyang, “Temporal-channel transformer for 3d lidar-based video object detection for autonomous driving,” T. Circuits and Systems for Video Technology , 2021
2021
Later among the works it cites.
D. Curto, A. Clapes, J. Selva, S. Smeureanu, J. C. S. J. Junior, D. Gallardo-Pujol, G. Guilera, D. Leiva, T. B. Moeslund, S. Escalera, and C. Palmero, “Dyadformer: A multi-modal transformer for long-range modeling of dyadic interactions,” in ICCV-W , 2021
2021
Later among the works it cites.
A. Prakash, K. Chitta, and A. Geiger, “Multi-modal fusion transformer for end-to-end autonomous driving,” in CVPR , 2021
2021
Later among the works it cites.
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in CVPR , 2022
2022
Closest in time.
T. Lin, Y. Wang, X. Liu, and X. Qiu, “A survey of transformers,” AI Open , 2022
2022
Closest in time.
Y. Xu, H. Wei, M. Lin, Y. Deng, K. Sheng, M. Zhang, F. Tang, W. Dong, F. Huang, and C. Xu, “Transformers in computational visual media: A survey,” Computational Visual Media , 2022
2022
Closest in time.
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM CSUR , 2022
2022
Closest in time.
Y. Yang, L. Jiao, X. Liu, F. Liu, S. Yang, Z. Feng, and X. Tang, “Transformers meet visual learning understanding: A comprehensive review,” arXiv , 2022
2022
Closest in time.
A. Shin, M. Ishii, and T. Narihira, “Perspectives and prospects on transformer architecture for cross-modal tasks with language and vision,” IJCV , 2022
2022
Closest in time.
P. Xu, X. Zhu, and D. A. Clifton, “Multimodal learning with transformers: a survey,” arXiv , 2022
2022
Closest in time.
L. Ruan and Q. Jin, “Survey: Transformer based video-language pre-training,” AI Open , 2022
2022
Closest in time.
C. Wei, H. Fan, S. Xie, C.-Y. Wu, A. Yuille, and C. Feichtenhofer, “Masked feature prediction for self-supervised visual pre-training,” in CVPR , 2022
2022
Closest in time.
Z. Tong, Y. Song, J. Wang, and L. Wang, “VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training,” in NeurIPS , 2022
2022
Closest in time.
R. Herzig, E. Ben-Avraham, K. Mangalam, A. Bar, G. Chechik, A. Rohrbach, T. Darrell, and A. Globerson, “Object-region video transformers,” in CVPR , 2022
2022
Closest in time.
Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong et al. , “Swin transformer v2: Scaling up capacity and resolution,” in CVPR , 2022
2022
Closest in time.
S. Yang, X. Wang, Y. Li, Y. Fang, J. Fang, W. Liu, X. Zhao, and Y. Shan, “Temporally efficient vision transformer for video instance segmentation,” in CVPR , 2022
2022
Closest in time.
T.-D. Truong, Q.-H. Bui, C. N. Duong, H.-S. Seo, S. L. Phung, X. Li, and K. Luu, “Direcformer: A directed attention in transformer approach to robust action recognition,” in CVPR , 2022
2022
Closest in time.
S. Yun, J. Kim, D. Han, H. Song, J.-W. Ha, and J. Shin, “Time is MattEr: Temporal self-supervision for video transformers,” in ICML , 2022
2022
Closest in time.
K. Ranasinghe, M. Naseer, S. Khan, F. S. Khan, and M. S. Ryoo, “Self-supervised video transformer,” in CVPR , 2022
2022
Closest in time.
Y. Li, C.-Y. Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, and C. Feichtenhofer, “Mvitv2: Improved multiscale vision transformers for classification and detection,” in CVPR , 2022
2022
Closest in time.
C.-Y. Wu, Y. Li, K. Mangalam, H. Fan, B. Xiong, J. Malik, and C. Feichtenhofer, “Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition,” in CVPR , 2022
2022
Closest in time.
J. Wang, G. Bertasius, D. Tran, and L. Torresani, “Long-short temporal contrastive learning of video transformers,” in CVPR , 2022
2022
Closest in time.
R. Wang, D. Chen, Z. Wu, Y. Chen, X. Dai, M. Liu, Y.-G. Jiang, L. Zhou, and L. Yuan, “Bevt: Bert pretraining of video transformers,” in CVPR , 2022
2022
Closest in time.
K. Li, Y. Wang, G. Peng, G. Song, Y. Liu, H. Li, and Y. Qiao, “Uniformer: Unified transformer for efficient spatial-temporal representation learning,” in ICLR , 2022
2022
Closest in time.
T. Meinhardt, A. Kirillov, L. Leal-Taixe, and C. Feichtenhofer, “Trackformer: Multi-object tracking with transformers,” in CVPR , 2022
2022
Closest in time.
S. Zheng, S. Chen, and Q. Jin, “Vrdformer: End-to-end video visual relation detection with transformers,” in CVPR , 2022
2022
Closest in time.
J. Yang, X. Dong, L. Liu, C. Zhang, J. Shen, and D. Yu, “Recurring the transformer for video action recognition,” in CVPR , 2022
2022
Closest in time.
S. Yan, X. Xiong, A. Arnab, Z. Lu, M. Zhang, C. Sun, and C. Schmid, “Multiview transformers for video recognition,” in CVPR , 2022
2022
Closest in time.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, “Scaling vision transformers,” CVPR , 2022
2022
Closest in time.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in CVPR , 2022
2022
Closest in time.
M. C. Schiappa, Y. S. Rawat, and M. Shah, “Self-supervised learning for videos: A survey,” arXiv , 2022
2022
Closest in time.
S. Guo, Z. Xiong, Y. Zhong, L. Wang, X. Guo, B. Han, and W. Huang, “Cross-architecture self-supervised video representation learning,” in CVPR , 2022
2022
Closest in time.
F. Xiao, K. Kundu, J. Tighe, and D. Modolo, “Hierarchical self-supervised representation learning for movie understanding,” in CVPR , 2022
2022
Closest in time.
R. Girdhar, M. Singh, N. Ravi, L. van der Maaten, A. Joulin, and I. Misra, “Omnivore: A single model for many visual modalities,” in CVPR , 2022
2022
Closest in time.
S. Buch, C. Eyzaguirre, A. Gaidon, J. Wu, L. Fei-Fei, and J. C. Niebles, “Revisiting the “video” in video-language understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022
2022
Closest in time.
L. Yuan, R. Qian, Y. Cui, B. Gong, F. Schroff, M.-H. Yang, H. Adam, and T. Liu, “Contextualized spatio-temporal contrastive learning with self-supervision,” in CVPR , 2022
2022
Closest in time.
T. Chen, Z. Zhang, Y. Cheng, A. Awadallah, and Z. Wang, “The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy,” in CVPR , 2022
2022
Closest in time.
C. Zhang, M. Zhang, S. Zhang, D. Jin, Q. Zhou, Z. Cai, H. Zhao, X. Liu, and Z. Liu, “Delving deep into the generalization of vision transformers under distribution shifts,” in CVPR , 2022
2022
Closest in time.
S. Paul and P.-Y. Chen, “Vision transformers are robust learners,” in AAI CAI , 2022
2022
Closest in time.
Y.-L. Sung, J. Cho, and M. Bansal, “Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks,” in CVPR , 2022
2022
Closest in time.
K. Lu, A. Grover, P. Abbeel, and I. Mordatch, “Frozen pretrained transformers as universal computation engines,” AAAI CAI , 2022
2022
Closest in time.
T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross, “Winoground: Probing vision and language models for visio-linguistic compositionality,” in CVPR , 2022
2022
Closest in time.
A. Bardes, J. Ponce, and Y. Lecun, “Vicreg: Variance-invariance-covariance regularization for self-supervised learning,” in ICLR , 2022
2022
Closest in time.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pvt v2: Improved baselines with pyramid vision transformer,” Computational Visual Media , vol. 8, 2022
2022
Closest in time.