Shao, Y. orcid.org/0009-0002-0475-7142, He, H. orcid.org/0009-0001-2554-2020, Li, S. orcid.org/0009-0008-5272-8657 et al. (11 more authors) (2025) EventVAD: Training-Free Event-Aware Video Anomaly Detection. In: Gurrin, C., Schoeffmann, K., Zhang, M., Rossetto, L., Rudinac, S., Dang-Nguyen, D.-T., Cheng, W.-H., Chen, P. and Benois-Pineau, J., (eds.) MM '25: Proceedings of the 33rd ACM International Conference on Multimedia. MM '25: The 33rd ACM International Conference on Multimedia, 27-31 Oct 2025, Dublin, Ireland. ACM, pp. 2586-2595. ISBN: 9798400720352.
Abstract
Video Anomaly Detection (VAD) focuses on identifying anomalies within videos. Supervised methods require an amount of in-domain training data and often struggle to generalize to unseen anomalies. In contrast, training-free methods leverage the intrinsic world knowledge of large language models (LLMs) to detect anomalies but face challenges in localizing fine-grained visual transitions and diverse events. Therefore, we propose EventVAD, an event-aware video anomaly detection framework that combines tailored dynamic graph architectures and multimodal LLMs to perform fine-grained temporal-event reasoning. Specifically, EventVAD first employs dynamic spatiotemporal graph modeling with time-decay constraints to capture event-aware video features. Then, it performs adaptive noise filtering and uses signal ratio thresholding to detect event boundaries via unsupervised statistical features. Finally, it utilizes a hierarchical prompting strategy to guide MLLMs in performing reasoning and making final decisions. We conducted extensive experiments on the UCF-Crime and XD-Violence datasets. The results demonstrate that EventVAD with a 7B MLLM achieves state-of-the-art (SOTA) in training-free settings, outperforming strong baselines that use 7B or larger MLLMs. The code is available at https://github.com/YihuaJerry/EventVAD.
Metadata
| Item Type: | Proceedings Paper |
|---|---|
| Authors/Creators: |
|
| Editors: |
|
| Copyright, Publisher and Additional Information: | © 2025 Copyright held by the owner/author(s). This work is licensed under a Creative Commons Attribution-NonCommercial International 4.0 License. (https://creativecommons.org/licenses/by-nc/4.0) |
| Keywords: | Machine Learning; Information and Computing Sciences; Artificial Intelligence; Computer Vision and Multimedia Computation; Bioengineering; Multimodal Large Language Models; Vision-Language Model; Video Understanding; Video Anomaly Detection |
| Dates: |
|
| Institution: | The University of Sheffield |
| Academic Units: | The University of Sheffield > Faculty of Engineering (Sheffield) > Department of Computer Science (Sheffield) |
| Date Deposited: | 03 Sep 2026 14:43 |
| Last Modified: | 03 Sep 2026 16:30 |
| Status: | Published |
| Publisher: | ACM |
| Refereed: | Yes |
| Identification Number: | 10.1145/3746027.3754500 |
| Related URLs: | |
| Open Archives Initiative ID (OAI ID): | oai:eprints.whiterose.ac.uk:245005 |
Download
Filename: 3746027.3754500.pdf
Licence: CC-BY-NC 4.0

CORE (COnnecting REpositories)
CORE (COnnecting REpositories)