Fetching the paper…
Reading the bibliography…
In this paper, we introduce Attention Prompt Tuning (APT) - a computationally efficient variant of prompt tuning for video-based applications such as action recognition.
RandAugment: Practical automated data augmentation with a reduced search space, 2019
Ekin D. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V. Le · 1909
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John Hertz · 1991
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
HMDB: A Large Video Database for Human Motion Recognition, 2011
H Kuehne, H Jhuang, E Garrote, T Poggio, and T Serre · 2011
Earlier work this paper cites.
UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild, 2012
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Differential recurrent neural networks for action recognition
Vivek Veeriah, Naifan Zhuang, and Guo-Jun Qi · 2015
Earlier work this paper cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
Limin Wang, Yu Qiao, and Xiaoou Tang · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Val Gool · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, and Trevor Darrell · 2017
Earlier work this paper cites.
The ”something something” video database for learning and evaluating visual common sense, 2017
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzyńska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic · 2017
Earlier work this paper cites.
The Kinetics Human Action Video Dataset, 2017
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Earlier work this paper cites.
Attention Is All You Need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Earlier work this paper cites.
Image Transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Earlier work this paper cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
Decoupled Weight Decay Regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Generative Pretraining From Pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever · 2020
Cited alongside, same era.
A comprehensive study of deep video action recognition
Yi Zhu, Xinyu Li, Chunhui Liu, Mohammadreza Zolfaghari, Yuanjun Xiong, Chongruo Wu, Zhi Zhang, Joseph Tighe, R Manmatha, and Mu Li · 2020
Cited alongside, same era.
Masked Autoencoders As Spatiotemporal Learners, 2022
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He · 2022
Later among the works it cites.
OmniMAE: Single Model Masked Pretraining on Images and Videos, 2022
Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2022
Later among the works it cites.
Ssast: Self-supervised audio spectrogram transformer
Yuan Gong, Cheng-I Lai, Yu-An Chung, and James Glass · 2022
Later among the works it cites.
A survey on vision transformer
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al · 2022
Later among the works it cites.
Towards a unified view of parameter-efficient transfer learning, 2022
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding Robustness of Transformers for Image Classification, 2021
Srinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li, Thomas Unterthiner, and Andreas Veit · 2021
Cited alongside, same era.
TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation, 2021
Jieneng Chen, Yongyi Lu, Qihang Yu, Xiangde Luo, Ehsan Adeli, Yan Wang, Le Lu, Alan L. Yuille, and Yuyin Zhou · 2021
Cited alongside, same era.
Masked Autoencoders Are Scalable Vision Learners, 2021
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning, 2021
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Cited alongside, same era.
Prefix-Tuning: Optimizing Continuous Prompts for Generation, 2021
Xiang Lisa Li and Percy Liang · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Spatiotemporal Contrastive Video Representation Learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui · 2021
Cited alongside, same era.
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim · 2022
Later among the works it cites.
P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang · 2022
Later among the works it cites.
Transformer-based sar image despeckling
Malsha V. Perera, Wele Gedara Chaminda Bandara, Jeya Maria Jose Valanarasu, and Vishal M. Patel · 2022
Later among the works it cites.
Spatio-Temporal Crop Aggregation for Video Representation Learning, 2022
Sepehr Sameni, Simon Jenni, and Paolo Favaro · 2022
Later among the works it cites.
Transformers in action recognition: A review on temporal modeling
Elham Shabaninia, Hossein Nezamabadi-pour, and Fatemeh Shafizadegan · 2022
Later among the works it cites.
Msp: Multi-stage prompting for making pre-trained language models better translators
Zhixing Tan, Xiangwen Zhang, Shuo Wang, and Yang Liu · 2022
Later among the works it cites.
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Later among the works it cites.
Dualprompt: Complementary prompting for rehearsal-free continual learning
Zifeng Wang, Zizhao Zhang, Sayna Ebrahimi, Ruoxi Sun, Han Zhang, Chen-Yu Lee, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, et al · 2022
Later among the works it cites.
StyleSwin: Transformer-Based GAN for High-Resolution Image Generation, 2022
Bowen Zhang, Shuyang Gu, Bo Zhang, Jianmin Bao, Dong Chen, Fang Wen, Yong Wang, and Baining Guo · 2022
Later among the works it cites.
E 2 vpt: An effective and efficient approach for visual prompt tuning, 2023
Cheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao, Wenguan Wang, Siyuan Qi, and Dongfang Liu · 2023
Later among the works it cites.
Fedperfix: Towards partial model personalization of vision transformers in federated learning
Guangyu Sun, Matias Mendieta, Jun Luo, Shandong Wu, and Chen Chen · 2023
Later among the works it cites.
Aim: Adapting image models for efficient video action recognition
Taojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang, Chen Chen, and Mu Li · 2023
Later among the works it cites.
Towards a unified view on visual parameter-efficient transfer learning, 2023
Bruce Yu, Jianlong Chang, Lingbo Liu, Qi Tian, and Chang Wen Chen · 2023
Later among the works it cites.