Fetching the paper…
Reading the bibliography…
Transferring knowledge from task-agnostic pre-trained deep models for downstream tasks is an important topic in computer vision research.
TDN: Temporal difference networks for efficient action recognition
Wang, L.; Tong, Z.; Ji, B.; and Wu, G. 2021 · 1904
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
DSANet: Dynamic Segment Aggregation Network for Video-Level Representation Learning
Wu, W.; Zhao, Y.; Xu, Y.; Tan, X.; He, D.; Zou, Z.; Ye, J.; Li, Y.; Yao, M.; Dong, Z.; et al. 2021b · 1911
Earlier work this paper cites.
Using discriminant analysis for multi-class classification: an experimental investigation
Li, T.; Zhu, S.; and Ogihara, M. 2006 · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Stm: Spatiotemporal and motion encoding for action recognition
Jiang, B.; Wang, M.; Gan, W.; Wu, W.; and Yan, J. 2019 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
Kuehne, H.; Jhuang, H.; Garrote, E.; Poggio, T.; and Serre, T. 2011 · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K.; Zamir, A. R.; and Shah, M. 2012 · 2012
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2014 · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; and Carlos Niebles, J. 2015 · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; et al. 2017 · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Qiu, Z.; Yao, T.; and Mei, T. 2017 · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C.; Shrivastava, A.; Singh, S.; and Gupta, A. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
A short note about kinetics-600
Carreira, J.; Noland, E.; Banki-Horvath, A.; Hillier, C.; and Zisserman, A. 2018 · 2018
Earlier work this paper cites.
A generative approach to zero-shot and few-shot action recognition
Mishra, A.; Verma, V. K.; Reddy, M. S. K.; Arulkumar, S.; Rai, P.; and Mittal, A. 2018 · 2018
Earlier work this paper cites.
A Closer Look at Spatiotemporal Convolutions for Action Recognition
Tran, D.; Wang, H.; Torresani, L.; Ray, J.; LeCun, Y.; and Paluri, M. 2018 · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Van den Oord, A.; Li, Y.; and Vinyals, O. 2018 · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Xie, S.; Sun, C.; Huang, J.; Tu, Z.; and Murphy, K. 2018 · 2018
Cited alongside, same era.
Slowfast networks for video recognition
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019 · 2019
Cited alongside, same era.
I know the relationships: Zero-shot action recognition via two-stream graph convolutional networks and knowledge graphs
Gao, J.; Zhang, T.; and Xu, C. 2019 · 2019
Cited alongside, same era.
Large-scale weakly-supervised pre-training for video action recognition
Ghadiyaram, D.; Tran, D.; and Mahajan, D. 2019 · 2019
Cited alongside, same era.
TSM: Temporal Shift Module for Efficient Video Understanding
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
Clipcap: Clip prefix for image captioning
Mokady, R.; Hertz, A.; and Bermano, A. H. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Ryoo, M. S.; Piergiovanni, A.; Arnab, A.; Dehghani, M.; and Angelova, A. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin, J.; Gan, C.; and Han, S. 2019 · 2019
Cited alongside, same era.
Video classification with channel-separated convolutional networks
Tran, D.; Wang, H.; Torresani, L.; and Feiszli, M. 2019 · 2019
Cited alongside, same era.
Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video Recognition
Wu, W.; He, D.; Tan, X.; Chen, S.; and Wen, S. 2019 · 2019
Cited alongside, same era.
Rethinking zero-shot video classification: End-to-end training for realistic applications
Brattoli, B.; Tighe, J.; Zhdanov, F.; Perona, P.; and Chalupka, K. 2020 · 2020
Cited alongside, same era.
X3D: Expanding Architectures for Efficient Video Recognition
Feichtenhofer, C. 2020 · 2020
Cited alongside, same era.
Listen to look: Action recognition by previewing audio
Gao, R.; Oh, T.-H.; Grauman, K.; and Torresani, L. 2020 · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020 · 2020
Cited alongside, same era.
Actionclip: A new paradigm for video action recognition
Wang, M.; Xing, J.; and Liu, Y. 2021 · 2021
Later among the works it cites.
T2vlad: global-local sequence alignment for text-video retrieval
Wang, X.; Zhu, L.; and Yang, Y. 2021 · 2021
Later among the works it cites.
Florence: A New Foundation Model for Computer Vision
Yuan, L.; Chen, D.; Chen, Y.-L.; Codella, N.; Dai, X.; Gao, J.; Hu, H.; Huang, X.; Li, B.; Li, C.; et al. 2021 · 2021
Later among the works it cites.
Co-training Transformer with Videos and Images Improves Action Recognition
Zhang, B.; Yu, J.; Fifty, C.; Han, W.; Dai, A. M.; Pang, R.; and Sha, F. 2021 · 2021
Later among the works it cites.
Learning to prompt for vision-language models
Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2021 · 2021
Later among the works it cites.
Mamico: Macro-to-micro semantic correspondence for self-supervised video representation learning
Fang, B.; Wu, W.; Liu, C.; Zhou, Y.; He, D.; and Wang, W. 2022 · 2022
Closest in time.
Prompting visual-language models for efficient video understanding
Ju, C.; Han, T.; Zheng, K.; Zhang, Y.; and Xie, W. 2022 · 2022
Closest in time.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022 · 2022
Closest in time.
Cross-modal Representation Learning for Zero-shot Action Recognition
Lin, C.-C.; Lin, K.; Wang, L.; Liu, Z.; and Li, L. 2022 · 2022
Closest in time.
Video swin transformer
Liu, Z.; Ning, J.; Cao, Y.; Wei, Y.; Zhang, Z.; Lin, S.; and Hu, H. 2022 · 2022
Closest in time.
CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning
Luo, H.; Ji, L.; Zhong, M.; Chen, Y.; Lei, W.; Duan, N.; and Li, T. 2022 · 2022
Closest in time.
Expanding language-image pretrained models for general video recognition
Ni, B.; Peng, H.; Chen, M.; Zhang, S.; Meng, G.; Fu, J.; Xiang, S.; and Ling, H. 2022 · 2022
Closest in time.
Multiview transformers for video recognition
Yan, S.; Xiong, X.; Arnab, A.; Lu, Z.; Zhang, M.; Sun, C.; and Schmid, C. 2022 · 2022
Closest in time.
Unified contrastive learning in image-text-label space
Yang, J.; Li, C.; Zhang, P.; Xiao, B.; Liu, C.; Yuan, L.; and Gao, J. 2022 · 2022
Closest in time.
CoCa: Contrastive Captioners are Image-Text Foundation Models
Yu, J.; Wang, Z.; Vasudevan, V.; Yeung, L.; Seyedhosseini, M.; and Wu, Y. 2022 · 2022
Closest in time.
Scaling vision transformers
Zhai, X.; Kolesnikov, A.; Houlsby, N.; and Beyer, L. 2022 · 2022
Closest in time.
CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
Zhao, S.; Zhu, L.; Wang, X.; and Yang, Y. 2022 · 2022
Closest in time.