Fetching the paper…
Reading the bibliography…
How can we enable models to comprehend video anomalies occurring over varying temporal scales and contexts? Traditional Video Anomaly Understanding (VAU) methods focus on frame-level anomaly prediction, often missing the interpretability of complex and diverse real-world anomalies.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Robust real-time unusual event detection using multiple fixed-location monitors
Amit Adam, Ehud Rivlin, Ilan Shimshoni, and Daviv Reinitz · 2008
Earlier work this paper cites.
Observe locally, infer globally: a space-time mrf for detecting abnormal activities with incremental updates
Jaechul Kim and Kristen Grauman · 2009
Earlier work this paper cites.
Abnormal crowd behavior detection using social force model
Ramin Mehran, Alexis Oyama, and Mubarak Shah · 2009
Earlier work this paper cites.
Online detection of unusual events in videos via dynamic sparse coding
Bin Zhao, Li Fei-Fei, and Eric P Xing · 2011
Earlier work this paper cites.
Anomaly detection and localization in crowded scenes
Weixin Li, Vijay Mahadevan, and Nuno Vasconcelos · 2013
Earlier work this paper cites.
Abnormal event detection at 150 fps in matlab
Cewu Lu, Jianping Shi, and Jiaya Jia · 2013
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Learning temporal regularity in video sequences
Mahmudul Hasan, Jonghyun Choi, Jan Neumann, Amit K Roy-Chowdhury, and Larry S Davis · 2016
Earlier work this paper cites.
Detecting anomalous events in videos by learning deep representations of appearance and motion
Dan Xu, Yan Yan, Elisa Ricci, and Nicu Sebe · 2017
Earlier work this paper cites.
Resound: Towards action recognition without representation bias
Yingwei Li, Yi Li, and Nuno Vasconcelos · 2018
Earlier work this paper cites.
Future frame prediction for anomaly detection–a new baseline
Wen Liu, Weixin Luo, Dongze Lian, and Shenghua Gao · 2018
Earlier work this paper cites.
Real-world anomaly detection in surveillance videos
Waqas Sultani, Chen Chen, and Mubarak Shah · 2018
Earlier work this paper cites.
Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection
Dong Gong, Lingqiao Liu, Vuong Le, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh, and Anton van den Hengel · 2019
Earlier work this paper cites.
Anomaly locality in video surveillance
Federico Landi, Cees GM Snoek, and Rita Cucchiara · 2019
Earlier work this paper cites.
Exploring background-bias for anomaly detection in surveillance videos
Kun Liu and Huadong Ma · 2019
Earlier work this paper cites.
Gods: Generalized one-class discriminative subspaces for anomaly detection
Jue Wang and Anoop Cherian · 2019
Earlier work this paper cites.
Graph convolutional label noise cleaner: Train a plug-and-play action classifier for anomaly detection
Jia-Xing Zhong, Nannan Li, Weijie Kong, Shan Liu, Thomas H Li, and Ge Li · 2019
Earlier work this paper cites.
Finegym: A hierarchical video dataset for fine-grained action understanding
Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin · 2020
Earlier work this paper cites.
Not only look, but also listen: Learning multimodal violence detection under weak supervision
Peng Wu, Jing Liu, Yujia Shi, Yujia Sun, Fangtao Shao, Zhaoyang Wu, and Zhiwei Yang · 2020
Earlier work this paper cites.
Mist: Multiple instance self-training framework for video anomaly detection
Jia-Chang Feng, Fa-Ting Hong, and Wei-Shi Zheng · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Coarse-fine networks for temporal activity detection in videos
Kumara Kahatapitiya and Michael S Ryoo · 2021
Cited alongside, same era.
A hybrid video anomaly detection framework via memory-augmented flow reconstruction and flow-guided frame prediction
Zhian Liu, Yongwei Nie, Chengjiang Long, Qing Zhang, and Guiqing Li · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Unbiased multiple instance learning for weakly supervised video anomaly detection
Hui Lv, Zhongqi Yue, Qianru Sun, Bin Luo, Zhen Cui, and Hanwang Zhang · 2023
Later among the works it cites.
Dyannet: A scene dynamicity guided self-trained video anomaly detection network
Kamalakar Vijay Thakare, Yash Raghuwanshi, Debi Prosad Dogra, Heeseung Choi, and Ig-Jae Kim · 2023
Later among the works it cites.
Exploring diffusion models for unsupervised video anomaly detection
Anil Osman Tur, Nicola Dall’Asen, Cigdem Beyan, and Elisa Ricci · 2023
Later among the works it cites.
Video event restoration based on keyframes for video anomaly detection
Zhiwei Yang, Jing Liu, Zhaoyang Wu, Peng Wu, and Xiaotao Liu · 2023
Later among the works it cites.
Towards surveillance video-and-language understanding: New dataset, baselines, and challenges, 2023
Tongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu, Chen Chen, Jian Jin, and Zhenzhen Jiao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weakly-supervised video anomaly detection with robust temporal feature magnitude learning
Yu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh, Johan W Verjans, and Gustavo Carneiro · 2021
Cited alongside, same era.
Self-training multi-sequence learning with transformer for weakly supervised video anomaly detection
Shuo Li, Fang Liu, and Licheng Jiao · 2022
Cited alongside, same era.
Fineaction: A fine-grained video dataset for temporal action localization
Yi Liu, Limin Wang, Yali Wang, Xiao Ma, and Yu Qiao · 2022
Cited alongside, same era.
Temporal global correlation network for end-to-end action proposal generation
Baiteng Ma, Shiwei Zhang, Changxin Gao, and Nong Sang · 2022
Cited alongside, same era.
Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu, Xiang Wang, Mingqian Tang, Changxin Gao, Rong Jin, and Nong Sang · 2022
Cited alongside, same era.
Review of action recognition based on multimodal data
SC Wang, Q Huang, YF Zhang, X Li, YQ Nie, and GC Luo · 2022
Cited alongside, same era.
Self-supervised sparse representation for video anomaly detection
Jhih-Ciang Wu, He-Yen Hsieh, Ding-Jie Chen, Chiou-Shann Fuh, and Tyng-Luh Liu · 2022
Cited alongside, same era.
Dual memory units with uncertainty regulation for weakly supervised video anomaly detection
Hang Zhou, Junqing Yu, and Wei Yang · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al · 2024
Closest in time.
Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al · 2024
Closest in time.
Adaptive sparse memory networks for efficient and robust video object segmentation
Jisheng Dang, Huicheng Zheng, Xiaohao Xu, Longguang Wang, Qingyong Hu, and Yulan Guo · 2024
Closest in time.
Uncovering what why and how: A comprehensive benchmark for causation understanding of video anomaly
Hang Du, Sicheng Zhang, Binzhu Xie, Guoshun Nan, Jiayang Zhang, Junrui Xu, Hangyu Liu, Sicong Leng, Jiangming Liu, Hehe Fan, et al · 2024
Closest in time.
Mmbench-video: A long-form multi-shot benchmark for holistic video understanding
Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, and Kai Chen · 2024
Closest in time.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al · 2024
Closest in time.
Video recap: Recursive captioning of hour-long videos
Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, and Gedas Bertasius · 2024
Closest in time.
Video anomaly detection and explanation via large language models
Hui Lv and Qianru Sun · 2024
Closest in time.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan · 2024
Closest in time.
Learning prompt-enhanced context features for weakly-supervised video anomaly detection
Yujiang Pu, Xiaoyu Wu, Lulu Yang, and Shengjin Wang · 2024
Closest in time.
Hawk: Learning to understand open-world video anomalies
Jiaqi Tang, Hao Lu, Ruizheng Wu, Xiaogang Xu, Ke Ma, Cheng Fang, Bin Guo, Jiangbo Lu, Qifeng Chen, and Ying-Cong Chen · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin · 2024
Closest in time.
Text prompt with normality guidance for weakly supervised video anomaly detection
Zhiwei Yang, Jing Liu, and Peng Wu · 2024
Closest in time.
Towards surveillance video-and-language understanding: New dataset baselines and challenges
Tongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu, Chen Chen, Jian Jin, and Zhenzhen Jiao · 2024
Closest in time.
Harnessing large language models for training-free video anomaly detection
Luca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang, and Elisa Ricci · 2024
Closest in time.