Fetching the paper…
Reading the bibliography…
Contrastive language-image pretraining (CLIP) has demonstrated remarkable success in various image tasks.
A new algorithm for the assignment problem
Dimitri P Bertsekas · 1981
Earlier work this paper cites.
Strike a pose: Tracking people by finding stylized poses
Deva Ramanan, David A Forsyth, and Andrew Zisserman · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Video ecommerce: Towards online video advertising
Cheng, Zhi-Qi and Liu, Yang and Wu, Xiao and Hua, Xian-Sheng · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Martin Arjovsky, Soumith Chintala, and Leon Bottou · 2017
Earlier work this paper cites.
Video2shop: Exact matching clothes in videos to online shopping images
Cheng, Zhi-Qi and Wu, Xiao and Liu, Yang and Hua, Xian-Sheng · 2017
Earlier work this paper cites.
Video ecommerce++: Toward large scale online video advertising
Cheng, Zhi-Qi and Wu, Xiao and Liu, Yang and Hua, Xian-Sheng · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Learning spatio-temporal features with 3d residual networks for action recognition
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2017
Earlier work this paper cites.
Tube convolutional neural network (t-cnn) for action detection in videos
Rui Hou, Chen Chen, and Mubarak Shah · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Rethinking the faster r-cnn architecture for temporal action localization
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A Ross, Jia Deng, and Rahul Sukthankar · 2018
Earlier work this paper cites.
Recurrent tubelet proposal and recognition networks for action detection
Dong Li, Zhaofan Qiu, Qi Dai, Ting Yao, and Tao Mei · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Fast online object tracking and segmentation: A unifying approach
Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip HS Torr · 2019
Earlier work this paper cites.
Edvr: Video restoration with enhanced deformable convolutional networks
Xintao Wang, Kelvin CK Chan, Ke Yu, Chao Dong, and Chen Change Loy · 2019
Earlier work this paper cites.
X3d: Expanding architectures for efficient video recognition
Christoph Feichtenhofer · 2020
Earlier work this paper cites.
Motionsqueeze: Neural motion feature learning for video understanding
Heeseung Kwon, Manjin Kim, Suha Kwak, and Minsu Cho · 2020
Earlier work this paper cites.
Something-else: Compositional action recognition with spatial-temporal interaction networks
Joanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu, Xiaolong Wang, and Trevor Darrell · 2020
Cited alongside, same era.
Weakly-supervised action localization by generative attention modeling
Baifeng Shi, Qi Dai, Yadong Mu, and Jingdong Wang · 2020
Cited alongside, same era.
Tdan: Temporally-deformable alignment network for video super-resolution
Yapeng Tian, Yulun Zhang, Yun Fu, and Chenliang Xu · 2020
Cited alongside, same era.
Single shot video object detector
Jiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou, Houqiang Li, and Tao Mei · 2020
Cited alongside, same era.
A dynamic frame selection framework for fast video recognition
Wu, Zuxuan and Li, Hengduo and Xiong, Caiming and Jiang, Yu-Gang and Davis, Larry Steven · 2020
Cited alongside, same era.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Basicvsr++: Improving video super-resolution with enhanced propagation and alignment
Kelvin CK Chan, Shangchen Zhou, Xiangyu Xu, and Chen Change Loy · 2022
Later among the works it cites.
Gsrformer: Grounded situation recognition transformer with alternate semantic attention refinement
Cheng, Zhi-Qi and Dai, Qi and Li, Siyao and Mitamura, Teruko and Hauptmann, Alexander · 2022
Later among the works it cites.
On the connection between local attention and dynamic depth-wise convolution
Qi Han, Zejia Fan, Qi Dai, Lei Sun, Ming-Ming Cheng, Jiaying Liu, and Jingdong Wang · 2022
Later among the works it cites.
Prompting visual-language models for efficient video understanding
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie · 2022
Later among the works it cites.
Uniformer: Unified transformer for efficient spatiotemporal representation learning
Kunchang Li, Yali Wang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Cited alongside, same era.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Cited alongside, same era.
Video super-resolution transformer
Jiezhang Cao, Yawei Li, Kai Zhang, and Luc Van Gool · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Cited alongside, same era.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Cited alongside, same era.
Clip-adapter: Better vision-language models with feature adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao · 2021
Cited alongside, same era.
Ta2n: Two-stage action alignment network for few-shot action recognition
Shuyuan Li, Huabin Liu, Rui Qian, Yuxi Li, John See, Mengjuan Fei, Xiaoyuan Yu, and Weiyao Lin · 2022
Later among the works it cites.
Improved multiscale vision transformers for classification and detection
Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer · 2022
Later among the works it cites.
Frozen clip models are efficient video learners
Ziyi Lin, Shijie Geng, Renrui Zhang, Peng Gao, Gerard de Melo, Xiaogang Wang, Jifeng Dai, Yu Qiao, and Hongsheng Li · 2022
Later among the works it cites.
Video swin transformer
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu · 2022
Later among the works it cites.
Expanding language-image pretrained models for general video recognition
Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling · 2022
Later among the works it cites.
St-adapter: Parameter-efficient image-to-video transfer learning for action recognition
Junting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao, and Hongsheng Li · 2022
Later among the works it cites.
Denseclip: Language-guided dense prediction with context-aware prompting
Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu · 2022
Later among the works it cites.
Rethinking alignment in video super-resolution transformers
Shuwei Shi, Jinjin Gu, Liangbin Xie, Xintao Wang, Yujiu Yang, and Chao Dong · 2022
Later among the works it cites.
Actionclip: A new paradigm for video action recognition
Mengmeng Wang, Jiazheng Xing, and Yong Liu · 2022
Later among the works it cites.
Multiview transformers for video recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid · 2022
Later among the works it cites.
Tip-adapter: Training-free clip-adapter for better vision-language modeling
Renrui Zhang, Rongyao Fang, Wei Zhang, Peng Gao, Kunchang Li, Jifeng Dai, Yu Qiao, and Hongsheng Li · 2022
Later among the works it cites.
Pointclip: Point cloud understanding by clip
Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li · 2022
Later among the works it cites.
Alignment-guided temporal attention for video action recognition
Yizhou Zhao, Zhenyang Li, Xun Guo, and Yan Lu · 2022
Later among the works it cites.
Omnivl: One foundation model for image-language and video-language tasks
Wang, Junke and Chen, Dongdong and Wu, Zuxuan and Luo, Chong and Zhou, Luowei and Zhao, Yucheng and Xie, Yujia and Liu, Ce and Jiang, Yu-Gang and Yuan, Lu · 2022
Later among the works it cites.
DAMO-StreamNet: Optimizing Streaming Perception in Autonomous Driving
He, Jun-Yan and Cheng, Zhi-Qi and Li, Chenyang and Xiang, Wangmeng and Chen, Binghui and Luo, Bin and Geng, Yifeng and Xie, Xuansong · 2023
Closest in time.
Resformer: Scaling vits with multi-resolution training
Rui Tian, Zuxuan Wu, Qi Dai, Han Hu, Yu Qiao, and Yu-Gang Jiang · 2023
Closest in time.
Svformer: Semi-supervised video transformer for action recognition
Zhen Xing, Qi Dai, Han Hu, Jingjing Chen, Zuxuan Wu, and Yu-Gang Jiang · 2023
Closest in time.
Aim: Adapting image models for efficient video action recognition
Taojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang, Chen Chen, and Mu Li · 2023
Closest in time.
Hivit: A simpler and more efficient design of hierarchical vision transformer
Xiaosong Zhang, Yunjie Tian, Lingxi Xie, Wei Huang, Qi Dai, Qixiang Ye, and Qi Tian · 2023
Closest in time.
Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight Optimization
Weng, Zejia and Yang, Xitong and Li, Ang and Wu, Zuxuan and Jiang, Yu-Gang · 2023
Closest in time.