Fetching the paper…
Reading the bibliography…
Recent Vision-Language Models (VLMs) \textit{e.g.} CLIP have made great progress in video recognition.
Ivan Laptev, ‘On space-time interest points’, IJCV
2005
Earlier work this paper cites.
Alexander Klaser, Marcin Marszałek, and Cordelia Schmid, ‘A spatio-temporal descriptor based on 3d-gradients’, in BMVC
2008
Earlier work this paper cites.
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre, ‘Hmdb: a large video database for human motion recognition’, in ICCV
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
Heng Wang, Alexander Kläser, Cordelia Schmid, and Cheng-Lin Liu, ‘Dense trajectories and motion boundary descriptors for action recognition’, IJCV
2013
Earlier work this paper cites.
R Christoph and Feichtenhofer Axel Pinz, ‘Spatiotemporal residual networks for video action recognition’, in NeurIPS
2016
Earlier work this paper cites.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, ‘Deep residual learning for image recognition’, in CVPR
2016
Earlier work this paper cites.
Joao Carreira and Andrew Zisserman, ‘Quo vadis, action recognition? a new model and the kinetics dataset’, in CVPR
2017
Earlier work this paper cites.
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al., ‘The” something something” video database for learning and evaluating visual common sense’, in ICCV
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, ‘Attention is all you need’, in NeurIPS
2017
Earlier work this paper cites.
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le, ‘Learning transferable architectures for scalable image recognition’, in CVPR
2018
Earlier work this paper cites.
Junyu Gao, Tianzhu Zhang, and Changsheng Xu, ‘I know the relationships: Zero-shot action recognition via two-stream graph convolutional networks and knowledge graphs’, in AAAI
2019
Earlier work this paper cites.
Biagio Brattoli, Joseph Tighe, Fedor Zhdanov, Pietro Perona, and Krzysztof Chalupka, ‘Rethinking zero-shot video classification: End-to-end training for realistic applications’, in CVPR
2020
Earlier work this paper cites.
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid, ‘Vivit: A video vision transformer’, in ICCV
2021
Earlier work this paper cites.
Gedas Bertasius, Heng Wang, and Lorenzo Torresani, ‘Is space-time attention all you need for video understanding?’, in ICML
2021
Earlier work this paper cites.
Gedas Bertasius, Heng Wang, and Lorenzo Torresani, ‘Is space-time attention all you need for video understanding?’, in ICML
2021
Cited alongside, same era.
Shizhe Chen and Dong Huang, ‘Elaborative rehearsal for zero-shot action recognition’, in ICCV
2021
Cited alongside, same era.
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al., ‘An image is worth 16x16 words: Transformers for image recognition at scale’, in ICLR
2021
Cited alongside, same era.
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig, ‘Scaling up visual and vision-language representation learning with noisy text supervision’, in ICML
2021
Cited alongside, same era.
Jingjia Huang, Yinan Li, Jiashi Feng, Xinglong Wu, Xiaoshuai Sun, and Rongrong Ji, ‘Clover: Towards a unified video-language alignment and fusion model’, in CVPR
2023
Later among the works it cites.
Kunchang Li, Yali Wang, Junhao Zhang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao, ‘Uniformer: Unifying convolution and self-attention for visual recognition’, T-PAMI
2023
Later among the works it cites.
2023
Later among the works it cites.
Shuyuan Tu, Qi Dai, Zuxuan Wu, Zhi-Qi Cheng, Han Hu, and Yu-Gang Jiang, ‘Implicit temporal modeling with learnable alignment for video recognition’, in ICCV
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mandela Patrick, Dylan Campbell, Yuki Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Joao F Henriques, ‘Keeping your eye on the ball: Trajectory attention in video transformers’, in NeurIPS
2021
Cited alongside, same era.
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., ‘Learning transferable visual models from natural language supervision’, in ICML
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie, ‘Prompting visual-language models for efficient video understanding’, in ECCV
2022
Cited alongside, same era.
Ziyi Lin, Shijie Geng, Renrui Zhang, Peng Gao, Gerard de Melo, Xiaogang Wang, Jifeng Dai, Yu Qiao, and Hongsheng Li, ‘Frozen clip models are efficient video learners’, in ECCV
2022
Cited alongside, same era.
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu, ‘Video swin transformer’, in CVPR
2022
Cited alongside, same era.
Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling, ‘Expanding language-image pretrained models for general video recognition’, in ECCV
2022
Cited alongside, same era.
Haohan Wang, Liang Liu, Wuhao Zhang, Jiangning Zhang, Zhenye Gan, Yabiao Wang, Chengjie Wang, and Haoqian Wang, ‘Iterative few-shot semantic segmentation from image label text’, in IJCAI
2023
Later among the works it cites.
Qiang Wang, Junlong Du, Ke Yan, and Shouhong Ding, ‘Seeing in flowing: Adapting clip for action recognition with motion prompts learning’, in ACM MM
2023
Later among the works it cites.
Syed Talal Wasim, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan, and Mubarak Shah, ‘Vita-clip: Video and text adaptive clip via multimodal prompting’, in CVPR
2023
Later among the works it cites.
Zejia Weng, Xitong Yang, Ang Li, Zuxuan Wu, and Yu-Gang Jiang, ‘Open-vclip: Transforming clip to an open-vocabulary video model via interpolated weight optimization’, in International Conference on Machine Learning
2023
Later among the works it cites.
Wenhao Wu, Zhun Sun, and Wanli Ouyang, ‘Revisiting classifier: Transferring vision-language models for video recognition’, in AAAI
2023
Later among the works it cites.
Wenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang, Yi Yang, and Wanli Ouyang, ‘Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models’, in CVPR
2023
Later among the works it cites.
Taojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang, Chen Chen, and Mu Li, ‘Aim: Adapting image models for efficient video action recognition’, in ICLR
2023
Later among the works it cites.
2024
Closest in time.
Mingyu Jin, Qinkai Yu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, Mengnan Du, et al., ‘The impact of reasoning step length on large language models’, in ACL
2024
Closest in time.
Ziqian Lu, Zhe-Ming Lu, Yunlong Yu, Zewei He, Hao Luo, and Yangming Zheng, ‘Learning multiple criteria calibration for generalized zero-shot learning’, Knowledge-Based Systems
2024
Closest in time.