Fetching the paper…
Reading the bibliography…
Vision Transformers have achieved impressive performance in video classification, while suffering from the quadratic complexity caused by the Softmax attention mechanism.
Object recognition from local scale-invariant features
David G Lowe · 1999
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Learning recursive filters for low-level vision via a hybrid neural network
Sifei Liu, Jinshan Pan, and Ming-Hsuan Yang · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Learning affinity via spatial propagation networks
Sifei Liu, Shalini De Mello, Jinwei Gu, Guangyu Zhong, Ming-Hsuan Yang, and Jan Kautz · 2017
Earlier work this paper cites.
Learnable pooling with context gating for video classification
Antoine Miech, Ivan Laptev, and Josef Sivic · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep learning using rectified linear units (relu)
Abien Fred Agarap · 2018
Earlier work this paper cites.
Action search: Spotting actions in videos and its application to temporal action localization
Humam Alwassel, Fabian Caba Heilbron, and Bernard Ghanem · 2018
Earlier work this paper cites.
A short note about kinetics-600
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman · 2018
Earlier work this paper cites.
Autoaugment: Learning augmentation policies from data
Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le · 2018
Earlier work this paper cites.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Earlier work this paper cites.
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun · 2018
Earlier work this paper cites.
Recurrent tubelet proposal and recognition networks for action detection
Dong Li, Zhaofan Qiu, Qi Dai, Ting Yao, and Tao Mei · 2018
Earlier work this paper cites.
Videolstm convolves, attends and flows for action recognition
Zhenyang Li, Kirill Gavrilyuk, Efstratios Gavves, Mihir Jain, and Cees GM Snoek · 2018
Earlier work this paper cites.
Generating wikipedia by summarizing long sequences
Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Earlier work this paper cites.
Temporal segment networks for action recognition in videos
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2018
Earlier work this paper cites.
Shift: A zero flop, zero parameter alternative to spatial convolutions
Bichen Wu, Alvin Wan, Xiangyu Yue, Peter Jin, Sicheng Zhao, Noah Golmant, Amir Gholaminejad, Joseph Gonzalez, and Kurt Keutzer · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Video action transformer network
Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman · 2019
Earlier work this paper cites.
Axial attention in multidimensional transformers
Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans · 2019
Earlier work this paper cites.
Ccnet: Criss-cross attention for semantic segmentation
Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu · 2019
Earlier work this paper cites.
Scsampler: Sampling salient clips from video for efficient action recognition
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2019
Earlier work this paper cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh · 2019
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Scott Wen-tau Yih, Sinong Wang, and Jie Tang · 2019
Earlier work this paper cites.
A theoretical analysis of contrastive unsupervised representation learning
Nikunj Saunshi, Orestis Plevrakis, Sanjeev Arora, Mikhail Khodak, and Hrishikesh Khandeparkar · 2019
Cited alongside, same era.
Video classification with channel-separated convolutional networks
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli · 2019
Cited alongside, same era.
Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition
Wenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen, and Shilei Wen · 2019
Cited alongside, same era.
Adaframe: Adaptive frame selection for fast video recognition
Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S Davis · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Cited alongside, same era.
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu · 2021
Later among the works it cites.
Tam: Temporal adaptive module for video recognition
Zhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian, and Tong Lu · 2021
Later among the works it cites.
Blind motion deblurring super-resolution: When dynamic spatio-temporal learning meets static image understanding
Wenjia Niu, Kaihao Zhang, Wenhan Luo, and Yiran Zhong · 2021
Later among the works it cites.
Keeping your eye on the ball: Trajectory attention in video transformers
Mandela Patrick, Dylan Campbell, Yuki Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and João F Henriques · 2021
Later among the works it cites.
Uncertainty class activation map (u-cam) using gradient certainty method
Badri Narayana Patro, Mayank Lunayach, and Vinay P Namboodiri · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, David Belanger, Lucy Colwell, et al · 2020
Cited alongside, same era.
X3d: Expanding architectures for efficient video recognition
Christoph Feichtenhofer · 2020
Cited alongside, same era.
Coot: Cooperative hierarchical transformer for video-text representation learning
Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox · 2020
Cited alongside, same era.
Pyramid constrained self-attention network for fast video salient object detection
Yuchao Gu, Lijuan Wang, Ziqin Wang, Yun Liu, Ming-Ming Cheng, and Shao-Ping Lu · 2020
Cited alongside, same era.
Learning channel-wise spatio-temporal representations for video salient object detection
Kan Huang, Ge Li, and Shan Liu · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong · 2021
Later among the works it cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh · 2021
Later among the works it cites.
An image is worth 16x16 words, what is a video worth?
Gilad Sharir, Asaf Noy, and Lihi Zelnik-Manor · 2021
Later among the works it cites.
Munet: Motion uncertainty-aware semi-supervised video object segmentation
Jiadai Sun, Yuxin Mao, Yuchao Dai, Yiran Zhong, and Jianyuan Wang · 2021
Later among the works it cites.
Contrastive learning, multi-view redundancy, and linear models
Christopher Tosh, Akshay Krishnamurthy, and Daniel Hsu · 2021
Later among the works it cites.
Efficient video transformers with spatial-temporal token selection
Junke Wang, Xitong Yang, Hengduo Li, Zuxuan Wu, and Yu-Gang Jiang · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Later among the works it cites.
Shifted chunk transformer for spatio-temporal representational learning
Xuefan Zha, Wentao Zhu, Lv Xun, Sen Yang, and Ji Liu · 2021
Later among the works it cites.
Token shift transformer for video classification
Hao Zhang, Yanbin Hao, and Chong-Wah Ngo · 2021
Later among the works it cites.
Vidtr: Video transformer without convolutions
Yanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe · 2021
Later among the works it cites.
Efficientvit: Enhanced linear attention for high-resolution low-computation visual recognition
Han Cai, Chuang Gan, and Song Han · 2022
Closest in time.
Implicit motion handling for video camouflaged object detection
Xuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong, Mehrtash Harandi, Tom Drummond, and Zongyuan Ge · 2022
Closest in time.
Deep laparoscopic stereo matching with transformers
Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Tom Drummond, Zhiyong Wang, and Zongyuan Ge · 2022
Closest in time.
Deep non-rigid structure-from-motion: A sequence-to-sequence translation perspective
Hui Deng, Tong Zhang, Yuchao Dai, Jiawei Shi, Yiran Zhong, and Hongdong Li · 2022
Closest in time.
Vision transformer with cross-attention by temporal shift for efficient action recognition
Ryota Hashiguchi and Toru Tamaki · 2022
Closest in time.
Neural architecture search on efficient transformers and beyond
Zexiang Liu, Dong Li, Kaiyue Lu, Zhen Qin, Weixuan Sun, Jiacheng Xu, and Yiran Zhong · 2022
Closest in time.
cosformer: Rethinking softmax in attention
Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong · 2022
Closest in time.
Locality matters: A locality-biased linear attention for automatic speech recognition
Jingyu Sun, Guiping Zhong, Dinghao Zhou, Baoxiang Li, and Yiran Zhong · 2022
Closest in time.
Weixuan Sun, Zhen Qin, Hui Deng, Jianyuan Wang, Yi Zhang, Kaihao Zhang, Nick Barnes, Stan Birchfield, Lingpeng Kong, and Yiran Zhong · 2022
Closest in time.
Inferring the class conditional response map for weakly supervised semantic segmentation
Weixuan Sun, Jing Zhang, and Nick Barnes · 2022
Closest in time.
Quadtree attention for vision transformers
Shitao Tang, Jiahui Zhang, Siyu Zhu, and Ping Tan · 2022
Closest in time.
Deformable video transformer
Jue Wang and Lorenzo Torresani · 2022
Closest in time.
An efficient spatio-temporal pyramid transformer for action detection
Yuetian Weng, Zizheng Pan, Mingfei Han, Xiaojun Chang, and Bohan Zhuang · 2022
Closest in time.
Flowformer: Linearizing transformers with conservation flows
Haixu Wu, Jialong Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long · 2022
Closest in time.
Linear complexity randomized self-attention mechanism
Lin Zheng, Chong Wang, and Lingpeng Kong · 2022
Closest in time.
Displacement-invariant cost computation for stereo matching
Yiran Zhong, Charles Loop, Wonmin Byeon, Stan Birchfield, Yuchao Dai, Kaihao Zhang, Alexey Kamenev, Thomas Breuel, Hongdong Li, and Jan Kautz · 2022
Closest in time.
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong · 2022
Closest in time.