A VLM Answer Is Not an Anomaly Score: Rank Compression Across Image and Video Anomaly Detection
Inpyo Song, Jangwon Lee
Abstract
Anomaly detection aims to identify observations that deviate from normal patterns. Recent work uses pretrained vision-language models (VLMs) for training-free image and video anomaly detection without task-specific retraining. Anomaly detection is commonly evaluated by how well anomaly scores rank anomalous images or video frames above normal ones. Generative VLMs, however, assign probabilities to possible answers and then decode a single answer. This decoding step can discard ordering information. We call this loss of ordering decoded-answer rank compression and study whether it materially affects anomaly detection performance. To isolate this effect, we compare two ways of scoring the same VLM output: one uses only the decoded answer, while the other computes a probability-weighted score over all possible answers. Across image and video anomaly detection benchmarks, VLMs, and answer scales, probability-weighted scoring consistently outperforms decoded-answer scoring, with mean gains ranging from 7.66 to 19.95 points on the primary benchmark metrics. Using answer probabilities only to break ties created by decoded-answer scoring recovers at least 95% of the average performance gap on every benchmark. When answer probabilities are available, how VLM answers are converted into anomaly scores is therefore part of the detector design, not merely an implementation detail.
Create a lesson
Related papers
PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image
Sneha Paul, Guile Wu, Bingbing Liu et al.
NormLift: From Lifted Features To Semantic Reliability In 3D Gaussian Splatting
Yihan Zang, Da Li, Dominik Engel et al.
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
Girish A. Koushik, Diptesh Kanojia, Helen Treharne
Copy What Is Seen, Generate What Is Not: Training-Free Anomaly-Aware Video Restoration
Zhida Qu, Shengchao Chen
Using OCR Heads to Verbalize Image Semantics
Sheridan Feucht, Benno Krojer, Sarah Wang et al.
DISTA-Net++: Rethinking Infrared Small Target Unmixing Beyond Sub-Pixel Separation
Mengze Xu, Zhu Liu, Weidong Sheng et al.