Fetching the paper…
Reading the bibliography…
Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Stm: Spatiotemporal and motion encoding for action recognition
Jiang, B.; Wang, M.; Gan, W.; Wu, W.; and Yan, J. 2019 · 2009
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
Kuehne, H.; Jhuang, H.; Garrote, E.; Poggio, T.; and Serre, T. 2011 · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K.; Zamir, A. R.; and Shah, M. 2012 · 2012
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Wang, L.; Xiong, Y.; Wang, Z.; Qiao, Y.; Lin, D.; Tang, X.; and Van Gool, L. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Goyal, R.; Ebrahimi Kahou, S.; Michalski, V.; Materzynska, J.; Westphal, S.; Kim, H.; Haenel, V.; Fruend, I.; Yianilos, P.; Mueller-Freitag, M.; et al. 2017 · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; et al. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Slowfast networks for video recognition
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding
Lin, J.; Gan, C.; and Han, S. 2019 · 2019
Earlier work this paper cites.
X3d: Expanding architectures for efficient video recognition
Feichtenhofer, C. 2020 · 2020
Cited alongside, same era.
Tea: Temporal excitation and aggregation for action recognition
Li, Y.; Ji, B.; Shi, X.; Zhang, J.; Kang, B.; and Wang, L. 2020 · 2020
Cited alongside, same era.
Vivit: A video vision transformer
Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lučić, M.; and Schmid, C. 2021 · 2021
Cited alongside, same era.
Is Space-Time Attention All You Need for Video Understanding?
Bertasius, G.; Wang, H.; and Torresani, L. 2021 · 2021
Cited alongside, same era.
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.; Parekh, Z.; Pham, H.; Le, Q. V.; Sung, Y.; Li, Z.; and Duerig, T. 2021 · 2021
Cited alongside, same era.
Keeping your eye on the ball: Trajectory attention in video transformers
Learning spatiotemporal and motion features in a unified 2d network for action recognition
Wang, M.; Xing, J.; Su, J.; Chen, J.; and Liu, Y. 2022 · 2022
Later among the works it cites.
Maple: Multi-modal prompt learning
Khattak, M. U.; Rasheed, H.; Maaz, M.; Khan, S.; and Khan, F. S. 2023 · 2023
Later among the works it cites.
Revisiting temporal modeling for clip-based image-to-video knowledge transferring
Liu, R.; Huang, J.; Li, G.; Feng, J.; Wu, X.; and Li, T. H. 2023 · 2023
Later among the works it cites.
Dual-path Adaptation from Image to Video Transformers
Park, J.; Lee, J.; and Sohn, K. 2023 · 2023
Later among the works it cites.
Implicit temporal modeling with learnable alignment for video recognition
Tu, S.; Dai, Q.; Wu, Z.; Cheng, Z.-Q.; Hu, H.; and Jiang, Y.-G. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Patrick, M.; Campbell, D.; Asano, Y.; Misra, I.; Metze, F.; Feichtenhofer, C.; Vedaldi, A.; and Henriques, J. F. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Florence: A New Foundation Model for Computer Vision
Yuan, L.; Chen, D.; Chen, Y.; Codella, N.; Dai, X.; Gao, J.; Hu, H.; Huang, X.; Li, B.; Li, C.; Liu, C.; Liu, M.; Liu, Z.; Lu, Y.; Shi, Y.; Wang, L.; Wang, J.; Xiao, B.; Xiao, Z.; Yang, J.; Zeng, M.; Zhou, L.; and Zhang, P. 2021 · 2021
Cited alongside, same era.
Prompting Visual-Language Models for Efficient Video Understanding
Ju, C.; Han, T.; Zheng, K.; Zhang, Y.; and Xie, W. 2022 · 2022
Cited alongside, same era.
Video swin transformer
Liu, Z.; Ning, J.; Cao, Y.; Wei, Y.; Zhang, Z.; Lin, S.; and Hu, H. 2022 · 2022
Cited alongside, same era.
Expanding Language-Image Pretrained Models for General Video Recognition
Ni, B.; Peng, H.; Chen, M.; Zhang, S.; Meng, G.; Fu, J.; Xiang, S.; and Ling, H. 2022 · 2022
Cited alongside, same era.
St-adapter: Parameter-efficient image-to-video transfer learning
Pan, J.; Lin, Z.; Zhu, X.; Shao, J.; and Li, H. 2022 · 2022
Cited alongside, same era.
Wang, M.; Xing, J.; Mei, J.; Liu, Y.; and Jiang, Y. 2023 · 2023
Later among the works it cites.
Vita-CLIP: Video and text adaptive CLIP via Multimodal Prompting
Wasim, S. T.; Naseer, M.; Khan, S.; Khan, F. S.; and Shah, M. 2023 · 2023
Later among the works it cites.
Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models
Wu, W.; Wang, X.; Luo, H.; Wang, J.; Yang, Y.; and Ouyang, W. 2023 · 2023
Later among the works it cites.
Multimodal Adaptation of CLIP for Few-Shot Action Recognition
Xing, J.; Wang, M.; Hou, X.; Dai, G.; Wang, J.; and Liu, Y. 2023 · 2023
Later among the works it cites.
Aim: Adapting image models for efficient video action recognition
Yang, T.; Zhu, Y.; Xie, Y.; Zhang, A.; Chen, C.; and Li, M. 2023 · 2023
Later among the works it cites.
Streaming Video Model
Zhao, Y.; Luo, C.; Tang, C.; Chen, D.; Codella, N.; and Zha, Z.-J. 2023 · 2023
Later among the works it cites.