Fetching the paper…
Reading the bibliography…
We introduce Video Transformer (VidTr) with separable-attention for video classification.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu · 2012
Earlier work this paper cites.
A dataset of 101 human action classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and M Shah · 2012
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Action recognition by learning deep multi-granular spatio-temporal video representation
Qing Li, Zhaofan Qiu, Ting Yao, Tao Mei, Yong Rui, and Jiebo Luo · 2016
Earlier work this paper cites.
VLAD3: Encoding Dynamics of Deep Features for Action Recognition
Yingwei Li, Weixin Li, Vijay Mahadevan, and Nuno Vasconcelos · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ivan Laptev, Ali Farhadi, and Abhinav Gupta · 2016
Earlier work this paper cites.
Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
ActionVLAD: Learning Spatio-Temporal Aggregation for Action Classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell · 2017
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Action recognition in video sequences using deep bi-directional lstm with cnn features
Amin Ullah, Jamil Ahmad, Khan Muhammad, Muhammad Sajjad, and Sung Wook Baik · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Earlier work this paper cites.
Multi-fiber networks for video recognition
Yunpeng Chen, Yannis Kalantidis, Jianshu Li, Shuicheng Yan, and Jiashi Feng · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Earlier work this paper cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Cited alongside, same era.
Temporal Relational Reasoning in Videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Cited alongside, same era.
End-to-end dense video captioning with masked transformer
Luowei Zhou, Yingbo Zhou, Jason J Corso, Richard Socher, and Caiming Xiong · 2018
Cited alongside, same era.
A short note on the kinetics-700 human action dataset
Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
TEA: Temporal Excitation and Aggregation for Action Recognition
Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang · 2020
Later among the works it cites.
Tea: Temporal excitation and aggregation for action recognition
Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang · 2020
Later among the works it cites.
Bridging text and video: A universal multimodal transformer for video-audio scene-aware dialog
Zekang Li, Zongjia Li, Jinchao Zhang, Yang Feng, Cheng Niu, and Jie Zhou · 2020
Later among the works it cites.
TEINet: Towards an Efficient Architecture for Video Recognition
Zhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Tong Lu · 2020
Later among the works it cites.
Temporal interlacing network
Hao Shao, Shengju Qian, and Yu Liu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Quanfu Fan, Chun-Fu Chen, Hilde Kuehne, Marco Pistoia, and David Cox · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
Video action transformer network
Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman · 2019
Cited alongside, same era.
STM: SpatioTemporal and Motion Encoding for Action Recognition
Boyuan Jiang, MengMeng Wang, Weihao Gan, Wei Wu, and Junjie Yan · 2019
Cited alongside, same era.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Cited alongside, same era.
AJ Piergiovanni, Anelia Angelova, and Michael S Ryoo · 2019
Cited alongside, same era.
Video classification with channel-separated convolutional networks
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli · 2019
Cited alongside, same era.
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
Later among the works it cites.
Video modeling with correlation networks
Heng Wang, Du Tran, Lorenzo Torresani, and Matt Feiszli · 2020
Later among the works it cites.
Temporal Pyramid Network for Action Recognition
Ceyuan Yang, Yinghao Xu, Jianping Shi, Bo Dai, and Bolei Zhou · 2020
Later among the works it cites.
Temporal pyramid network for action recognition
Ceyuan Yang, Yinghao Xu, Jianping Shi, Bo Dai, and Bolei Zhou · 2020
Later among the works it cites.
Transpose: Towards explainable human pose estimation by transformer
Sen Yang, Zhibin Quan, Mu Nie, and Wankou Yang · 2020
Later among the works it cites.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Closest in time.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
An empirical study of training self-supervised visual transformers
Xinlei Chen, Saining Xie, and Kaiming He · 2021
Closest in time.
Sstvos: Sparse spatiotemporal transformers for video object segmentation
Brendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi, and Graham W Taylor · 2021
Closest in time.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Closest in time.
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu · 2021
Closest in time.
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann · 2021
Closest in time.
Keeping your eye on the ball: Trajectory attention in video transformers
Mandela Patrick, Dylan Campbell, Yuki M Asano, Ishan Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, Jo Henriques, et al · 2021
Closest in time.
Cvt: Introducing convolutions to vision transformers
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Closest in time.