Fetching the paper…
Reading the bibliography…
Recently, large-scale pre-trained vision-language models (e.g., CLIP), have garnered significant attention thanks to their powerful representative capabilities.
L. Wang, Z. Tong, B. Ji, and G. Wu, “Tdn: Temporal difference networks for efficient action recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 1895–1904
1904
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag et al. , “The” something something” video database for learning and evaluating visual common sense,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5842–5850
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in International conference on machine learning . PMLR, 2019, pp. 2790–2799
2019
Earlier work this paper cites.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 6202–6211
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Z. Jiang, F. F. Xu, J. Araki, and G. Neubig, “How can we know what language models know?” Transactions of the Association for Computational Linguistics , vol. 8, pp. 423–438, 2020
2020
Earlier work this paper cites.
J. O. Zhang, A. Sax, A. Zamir, L. Guibas, and J. Malik, “Side-tuning: a baseline for network adaptation via additive side networks,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16 . Springer, 2020, pp. 698–714
2020
Earlier work this paper cites.
Y. Li, B. Ji, X. Shi, J. Zhang, B. Kang, and L. Wang, “Tea: Temporal excitation and aggregation for action recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 909–918
2020
Earlier work this paper cites.
M. Ryoo, A. Piergiovanni, A. Arnab, M. Dehghani, and A. Angelova, “Tokenlearner: Adaptive space-time tokenization for videos,” Advances in neural information processing systems , vol. 34, pp. 12 786–12 797, 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Wang, Q. She, and A. Smolic, “Action-net: Multipath excitation for action recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 13 214–13 223
2021
Cited alongside, same era.
H. Zhang, Y. Hao, and C.-W. Ngo, “Token shift transformer for video classification,” in Proceedings of the 29th ACM International Conference on Multimedia , 2021, pp. 917–925
2021
Cited alongside, same era.
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid, “Vivit: A video vision transformer,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 6836–6846
2021
Cited alongside, same era.
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 6824–6835
2021
Cited alongside, same era.
Z. Qing, S. Zhang, Z. Huang, Y. Zhang, C. Gao, D. Zhao, and N. Sang, “Disentangling spatial and temporal learning for efficient image-to-video transfer learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 13 934–13 944
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Wu, Z. Sun, and W. Ouyang, “Revisiting classifier: Transferring vision-language models for video recognition,” in Proceedings of the AAAI conference on artificial intelligence , vol. 37, no. 3, 2023, pp. 2847–2855
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Shu, B. Xu, L. Zhang, and J. Tang, “Multi-granularity anchor-contrastive representation learning for semi-supervised skeleton-based action recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 6, pp. 7559–7576, 2022
2022
Cited alongside, same era.
Z. Lin, S. Geng, R. Zhang, P. Gao, G. De Melo, X. Wang, J. Dai, Y. Qiao, and H. Li, “Frozen clip models are efficient video learners,” in European Conference on Computer Vision . Springer, 2022, pp. 388–404
2022
Cited alongside, same era.
J. Pan, Z. Lin, X. Zhu, J. Shao, and H. Li, “St-adapter: Parameter-efficient image-to-video transfer learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 26 462–26 477, 2022
2022
Cited alongside, same era.
Y.-L. Sung, J. Cho, and M. Bansal, “Lst: Ladder side-tuning for parameter and memory efficient transfer learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 12 991–13 005, 2022
2022
Cited alongside, same era.
J. Zhou, Z. Fu, Q. Huang, Q. Liu, and Y. Wang, “Lgnet: A local-global network for action recognition and beyond,” IEEE Transactions on Multimedia , vol. 25, pp. 5192–5205, 2022
2022
Cited alongside, same era.
M. Jia, L. Tang, B.-C. Chen, C. Cardie, S. Belongie, B. Hariharan, and S.-N. Lim, “Visual prompt tuning,” in European conference on computer vision . Springer, 2022, pp. 709–727
2022
Cited alongside, same era.
R. Herzig, E. Ben-Avraham, K. Mangalam, A. Bar, G. Chechik, A. Rohrbach, T. Darrell, and A. Globerson, “Object-region video transformers,” in Proceedings of the ieee/cvf conference on computer vision and pattern recognition , 2022, pp. 3148–3159
2022
Cited alongside, same era.
Y. Liu, P. Xiong, L. Xu, S. Cao, and Q. Jin, “Ts2-net: Token shift and selection transformer for text-video retrieval,” in European conference on computer vision . Springer, 2022, pp. 319–335
2022
Cited alongside, same era.
M. Wang, J. Xing, J. Mei, Y. Liu, and Y. Jiang, “Actionclip: Adapting language-image pretrained models for video action recognition,” IEEE Transactions on Neural Networks and Learning Systems , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Park, J. Lee, K. Sohn et al. , “Dual-path adaptation from image to video transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2203–2213
2023
Later among the works it cites.
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, L. Wang, and Y. Qiao, “Uniformerv2: Unlocking the potential of image vits for video understanding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 1632–1643
2023
Later among the works it cites.
S. Tu, Q. Dai, Z. Wu, Z.-Q. Cheng, H. Hu, and Y.-G. Jiang, “Implicit temporal modeling with learnable alignment for video recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 19 936–19 947
2023
Later among the works it cites.
Y. Zhao, C. Luo, C. Tang, D. Chen, N. Codella, and Z.-J. Zha, “Streaming video model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 14 602–14 612
2023
Later among the works it cites.
R. Liu, J. Huang, G. Li, J. Feng, X. Wu, and T. H. Li, “Revisiting temporal modeling for clip-based image-to-video knowledge transferring,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 6555–6564
2023
Later among the works it cites.
W. Wu, X. Wang, H. Luo, J. Wang, Y. Yang, and W. Ouyang, “Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2023, pp. 6620–6630
2023
Later among the works it cites.
H. Wang, F. Liu, L. Jiao, J. Wang, Z. Hao, S. Li, L. Li, P. Chen, and X. Liu, “Vilt-clip: Video and language tuning clip with multimodal prompt learning and scenario-guided optimization,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 6, 2024, pp. 5390–5400
2024
Closest in time.
X. Liu, P. Zhang, C. Yu, X. Qian, X. Yang, and H. Lu, “A video is worth three views: Trigeminal transformers for video-based person re-identification,” IEEE Transactions on Intelligent Transportation Systems , 2024
2024
Closest in time.
M. Wang, J. Xing, B. Jiang, J. Chen, J. Mei, X. Zuo, G. Dai, J. Wang, and Y. Liu, “A multimodal, multi-task adapting framework for video action recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 6, 2024, pp. 5517–5525
2024
Closest in time.
C. Yu, X. Liu, Y. Wang, P. Zhang, and H. Lu, “Tf-clip: Learning text-free clip for video-based person re-identification,” in Proceedings of the AAAI conference on artificial intelligence , vol. 38, no. 7, 2024, pp. 6764–6772
2024
Closest in time.
G. Dai, X. Shu, W. Wu, R. Yan, and J. Zhang, “Gpt4ego: unleashing the potential of pre-trained models for zero-shot egocentric action recognition,” IEEE Transactions on Multimedia , 2024
2024
Closest in time.