Fetching the paper…
Reading the bibliography…
Video Anomaly Detection~(VAD) focuses on identifying anomalies within videos.
Position enhanced mention graph attention network for dialogue relation extraction. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1985–1989
Xinwei Long, Shuzi Niu, and Yucheng Li. 2021 · 1989
Earlier work this paper cites.
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. 2008 · 2008
Earlier work this paper cites.
Detecting outliers: Do not use standard deviation around the mean, use absolute deviation around the median
Christophe Leys, Christophe Ley, Olivier Klein, Philippe Bernard, and Laurent Licata. 2013 · 2013
Earlier work this paper cites.
Abnormal event detection at 150 fps in matlab. In ICCV
Cewu Lu, Jianping Shi, and Jiaya Jia. 2013 · 2013
Earlier work this paper cites.
Learning temporal regularity in video sequences. In CVPR
Mahmudul Hasan, Jonghyun Choi, Jan Neumann, Amit K Roy-Chowdhury, and Larry S Davis. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Subspace support vector data description. In ICPR
Fahad Sohrab, Jenni Raitoharju, Moncef Gabbouj, and Alexandros Iosifidis. 2018 · 2018
Earlier work this paper cites.
Real-world anomaly detection in surveillance videos. In CVPR
Waqas Sultani, Chen Chen, and Mubarak Shah. 2018 · 2018
Earlier work this paper cites.
Graph Attention Networks. In International Conference on Learning Representations
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018 · 2018
Earlier work this paper cites.
Gods: Generalized one-class discriminative subspaces for anomaly detection. In ICCV
Jue Wang and Anoop Cherian. 2019 · 2019
Earlier work this paper cites.
A survey of single-scene video anomaly detection
Bharathkumar Ramachandra, Michael J Jones, and Ranga Raju Vatsavai. 2020 · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16 . Springer, 402–419
Zachary Teed and Jia Deng. 2020 · 2020
Earlier work this paper cites.
Not only look, but also listen: Learning multimodal violence detection under weak supervision. In ECCV
Peng Wu, Jing Liu, Yujia Shi, Yujia Sun, Fangtao Shao, Zhaoyang Wu, and Zhiwei Yang. 2020 · 2020
Earlier work this paper cites.
Claws: Clustering assisted weakly supervised learning with normalcy suppression for anomalous event detection. In ECCV
Muhammad Zaigham Zaheer, Arif Mahmood, Marcella Astrid, and Seung-Ik Lee. 2020 · 2020
Earlier work this paper cites.
A comprehensive review on deep learning-based methods for video anomaly detection
Rashmiranjan Nayak, Umesh Chandra Pati, and Santos Kumar Das. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PmLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Weakly-supervised video anomaly detection with robust temporal feature magnitude learning. In ICCV
Yu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh, Johan W Verjans, and Gustavo Carneiro. 2021 · 2021
Earlier work this paper cites.
Learning causal temporal relation and feature discrimination for anomaly detection
Peng Wu and Jing Liu. 2021 · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
Self-supervised Sparse Representation for Video Anomaly Detection. In ECCV
Jhih-Ciang Wu, He-Yen Hsieh, Ding-Jie Chen, Chiou-Shann Fuh, and Tyng-Luh Liu. 2022 · 2022
Earlier work this paper cites.
Generative cooperative learning for unsupervised video anomaly detection. In CVPR
M Zaigham Zaheer, Arif Mahmood, M Haris Khan, Mattia Segu, Fisher Yu, and Seung-Ik Lee. 2022 · 2022
Earlier work this paper cites.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. 2023b · 2023
Earlier work this paper cites.
Imagebind: One embedding space to bind them all. In CVPR
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023 · 2023
Cited alongside, same era.
Unified language-vision pretraining in llm with dynamic discrete visual tokenization
Yang Jin, Kun Xu, Liwei Chen, Chao Liao, Jianchao Tan, Quzhe Huang, Bin Chen, Chenyi Lei, An Liu, Chengru Song, et al · 2023
Cited alongside, same era.
CLIP-TSA: CLIP-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection. In ICIP
Hyekang Kevin Joo, Khoa Vo, Kashu Yamazaki, and Ngan Le. 2023 · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Cited alongside, same era.
Video-llava: Learning united visual representation by alignment before projection
Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22195–22206
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al · 2024
Later among the works it cites.
St-llm: Large language models are effective temporal learners. In European Conference on Computer Vision . Springer, 1–18
Ruyang Liu, Chen Li, Haoran Tang, Yixiao Ge, Ying Shan, and Ge Li. 2024 · 2024
Later among the works it cites.
Trust in internal or external knowledge? generative multi-modal entity linking with knowledge retriever. In Findings of the Association for Computational Linguistics ACL 2024 . 7559–7569
Xinwei Long, Jiali Zeng, Fandong Meng, Jie Zhou, and Bowen Zhou. 2024b · 2024
Later among the works it cites.
GWQ: Gradient-Aware Weight Quantization for Large Language Models
Yihua Shao, Siyu Liang, Zijian Ling, Minxi Yan, Haiyang Liu, Siyu Chen, Ziyang Yan, Chenyu Zhang, Haotong Qin, Michele Magno, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. 2023 · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023a · 2023
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023b · 2023
Cited alongside, same era.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2023 · 2023
Cited alongside, same era.
A critical analysis of NeRF-based 3D reconstruction
Fabio Remondino, Ali Karami, Ziyang Yan, Gabriele Mazzacca, Simone Rigon, and Rongjun Qin. 2023 · 2023
Cited alongside, same era.
Pandagpt: One model to instruction-follow them all
Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai. 2023 · 2023
Cited alongside, same era.
Eva-clip: Improved training techniques for clip at scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. 2023 · 2023
Cited alongside, same era.
RareAnom: A Benchmark Video Dataset for Rare Type Anomalies
Kamalakar Vijay Thakare, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, and Ig-Jae Kim. 2023a · 2023
Cited alongside, same era.
AccidentBlip: Agent of Accident Warning based on MA-former
Yihua Shao, Yeling Xu, Xinwei Long, Siyu Chen, Ziyang Yan, Yang Yang, Haoting Liu, Yan Wang, Hao Tang, and Zhen Lei. 2024b · 2024
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models
Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Song XiXuan, et al · 2024
Later among the works it cites.
Renderworld: World model with self-supervised 3d label
Ziyang Yan, Wenzhen Dong, Yihua Shao, Yuhang Lu, Liu Haiyang, Jingwen Liu, Haozhe Wang, Zhe Wang, Yan Wang, Fabio Remondino, et al · 2024
Later among the works it cites.
3dsceneeditor: Controllable 3d scene editing with gaussian splatting
Ziyang Yan, Lei Li, Yihua Shao, Siyu Chen, Zongkai Wu, Jenq-Neng Hwang, Hao Zhao, and Fabio Remondino. 2024b · 2024
Later among the works it cites.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition . 13040–13051
Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Anwen Hu, Haowei Liu, Qi Qian, Ji Zhang, and Fei Huang. 2024 · 2024
Later among the works it cites.
Harnessing large language models for training-free video anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18527–18536
Luca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang, and Elisa Ricci. 2024 · 2024
Later among the works it cites.
M2M-TAG: Training-Free Many-to-Many Token Aggregation for Vision Transformer Acceleration. In Workshop on Machine Learning and Compression, NeurIPS 2024
Fanhu Zeng and Deli Yu. 2024 · 2024
Later among the works it cites.
Modalprompt: Dual-modality guided prompt for continual learning of large multimodal models
Fanhu Zeng, Fei Zhu, Haiyang Guo, Xu-Yao Zhang, and Cheng-Lin Liu. 2024 · 2024
Later among the works it cites.
MCANet: Multimodal Caption Aware Training-Free Video Anomaly Detection via Large Language Model. In International Conference on Pattern Recognition . Springer, 362–379
Prabhu Prasad Dev, Raju Hazari, and Pranesh Das. 2025 · 2025
Closest in time.
GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-Experts
Minwen Liao, Hao Bo Dong, Xinyi Wang, Ziyang Yan, and Yihua Shao. 2025 · 2025
Closest in time.
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 24723–24731
Xinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang, Biqing Qi, and Bowen Zhou. 2025 · 2025
Closest in time.
TR-DQ: Time-Rotation Diffusion Quantization
Yihua Shao, Deyang Lin, Fanhu Zeng, Minxi Yan, Muyang Zhang, Siyu Chen, Yuxuan Fan, Ziyang Yan, Haozhe Wang, Jingcai Guo, et al · 2025
Closest in time.
In-Context Meta LoRA Generation
Yihua Shao, Minxi Yan, Yang Liu, Siyu Chen, Wenjie Chen, Xinwei Long, Ziyang Yan, Lei Li, Chenyu Zhang, Nicu Sebe, et al · 2025
Closest in time.
Unifying Appearance Codes and Bilateral Grids for Driving Scene Gaussian Splatting
Nan Wang, Yuantao Chen, Lixing Xiao, Weiqing Xiao, Bohan Li, Zhaoxi Chen, Chongjie Ye, Shaocong Xu, Saining Zhang, Ziyang Yan, et al · 2025
Closest in time.
Learning-Based 3D Reconstruction Methods for Non-Collaborative Surfaces—A Metrological Evaluation
Ziyang Yan, Nazanin Padkan, Paweł Trybała, Elisa Mariarosaria Farella, and Fabio Remondino. 2025 · 2025
Closest in time.
Fanhu Zeng, Zhen Cheng, Fei Zhu, and Xu-Yao Zhang. 2025a · 2025
Closest in time.
Fanhu Zeng, Haiyang Guo, Fei Zhu, Li Shen, and Hao Tang. 2025c · 2025
Closest in time.
MambaIC: State Space Models for High-Performance Learned Image Compression
Fanhu Zeng, Hao Tang, Yihua Shao, Siyu Chen, Ling Shao, and Yan Wang. 2025d · 2025
Closest in time.