Fetching the paper…
Reading the bibliography…
The recent contrastive language-image pre-training (CLIP) model has shown great success in a wide range of image-level tasks, revealing remarkable ability for learning powerful visual representations with rich semantics.
Support vector method for novelty detection
Schölkopf, B.; Williamson, R. C.; Smola, A.; Shawe-Taylor, J.; and Platt, J. 1999 · 1999
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; and Paluri, M. 2015 · 2015
Earlier work this paper cites.
Learning temporal regularity in video sequences
Hasan, M.; Choi, J.; Neumann, J.; Roy-Chowdhury, A. K.; and Davis, L. S. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Real-world anomaly detection in surveillance videos
Sultani, W.; Chen, C.; and Shah, M. 2018 · 2018
Earlier work this paper cites.
Graph convolutional label noise cleaner: Train a plug-and-play action classifier for anomaly detection
Zhong, J.-X.; Li, N.; Kong, W.; Liu, S.; Li, T. H.; and Li, G. 2019 · 2019
Earlier work this paper cites.
Not only look, but also listen: Learning multimodal violence detection under weak supervision
Wu, P.; Liu, J.; Shi, Y.; Sun, Y.; Shao, F.; Wu, Z.; and Yang, Z. 2020 · 2020
Earlier work this paper cites.
Claws: Clustering assisted weakly supervised learning with normalcy suppression for anomalous event detection
Zaheer, M. Z.; Mahmood, A.; Astrid, M.; and Lee, S.-I. 2020 · 2020
Earlier work this paper cites.
Mist: Multiple instance self-training framework for video anomaly detection
Feng, J.-C.; Hong, F.-T.; and Zheng, W.-S. 2021 · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.-T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.-H.; Li, Z.; and Duerig, T. 2021 · 2021
Earlier work this paper cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W.; Son, B.; and Kim, I. 2021 · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Cited alongside, same era.
Clipcap: Clip prefix for image captioning
Mokady, R.; Hertz, A.; and Bermano, A. H. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Weakly-supervised video anomaly detection with robust temporal feature magnitude learning
Tian, Y.; Pang, G.; Chen, Y.; Singh, R.; Verjans, J. W.; and Carneiro, G. 2021 · 2021
Cited alongside, same era.
Actionclip: A new paradigm for video action recognition
Wang, M.; Xing, J.; and Liu, Y. 2021 · 2021
Cited alongside, same era.
CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning
Luo, H.; Ji, L.; Zhong, M.; Chen, Y.; Lei, W.; Duan, N.; and Li, T. 2022 · 2022
Later among the works it cites.
Zero-shot temporal action detection via vision-language prompting
Nag, S.; Zhu, X.; Song, Y.-Z.; and Xiang, T. 2022 · 2022
Later among the works it cites.
Expanding language-image pretrained models for general video recognition
Ni, B.; Peng, H.; Chen, M.; Zhang, S.; Meng, G.; Fu, J.; Xiang, S.; and Ling, H. 2022 · 2022
Later among the works it cites.
Denseclip: Language-guided dense prediction with context-aware prompting
Rao, Y.; Zhao, W.; Chen, G.; Tang, Y.; Zhu, Z.; Huang, G.; Zhou, J.; and Lu, J. 2022 · 2022
Later among the works it cites.
Weakly supervised audio-visual violence detection
Wu, P.; Liu, X.; and Liu, J. 2022 · 2022
Later among the works it cites.
Detecting twenty-thousand classes using image-level supervision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simvlm: Simple visual language model pretraining with weak supervision
Wang, Z.; Yu, J.; Yu, A. W.; Dai, Z.; Tsvetkov, Y.; and Cao, Y. 2021 · 2021
Cited alongside, same era.
Weakly-supervised spatio-temporal anomaly detection in surveillance video
Wu, J.; Zhang, W.; Li, G.; Wu, W.; Tan, X.; Li, Y.; Ding, E.; and Lin, L. 2021 · 2021
Cited alongside, same era.
Learning causal temporal relation and feature discrimination for anomaly detection
Wu, P.; and Liu, J. 2021 · 2021
Cited alongside, same era.
Weakly Supervised Video Anomaly Detection via Self-Guided Temporal Discriminative Transformer
Huang, C.; Liu, C.; Wen, J.; Wu, L.; Xu, Y.; Jiang, Q.; and Wang, Y. 2022 · 2022
Cited alongside, same era.
Prompting visual-language models for efficient video understanding
Ju, C.; Han, T.; Zheng, K.; Zhang, Y.; and Xie, W. 2022 · 2022
Cited alongside, same era.
Self-training multi-sequence learning with transformer for weakly supervised video anomaly detection
Li, S.; Liu, F.; and Jiao, L. 2022 · 2022
Cited alongside, same era.
Frozen clip models are efficient video learners
Lin, Z.; Geng, S.; Zhang, R.; Gao, P.; de Melo, G.; Wang, X.; Dai, J.; Qiao, Y.; and Li, H. 2022 · 2022
Cited alongside, same era.
Zhou, X.; Girdhar, R.; Joulin, A.; Krähenbühl, P.; and Misra, I. 2022b · 2022
Later among the works it cites.
Clip-tsa: Clip-assisted temporal self-attention for weakly-supervised video anomaly detection
Joo, H. K.; Vo, K.; Yamazaki, K.; and Le, N. 2023 · 2023
Closest in time.
Unbiased Multiple Instance Learning for Weakly Supervised Video Anomaly Detection
Lv, H.; Yue, Z.; Sun, Q.; Luo, B.; Cui, Z.; and Zhang, H. 2023 · 2023
Closest in time.
Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model
Wu, P.; Liu, J.; He, X.; Peng, Y.; Wang, P.; and Zhang, Y. 2023 · 2023
Closest in time.
Turning a CLIP Model into a Scene Text Detector
Yu, W.; Liu, Y.; Hua, W.; Jiang, D.; Ren, B.; and Bai, X. 2023 · 2023
Closest in time.
Dual Memory Units with Uncertainty Regulation for Weakly Supervised Video Anomaly Detection
Zhou, H.; Yu, J.; and Yang, W. 2023 · 2023
Closest in time.