VCF-CLIP: Visual Context-Driven Fine-Grained Prompt Learning for Zero-Shot Anomaly Detection.

Fu, Kaiwen; Qi, Fei; Chang, Chengyuan; Chen, Jiahui; Zhao, Zhifu; Wang, Xiaotian; Liu, Kun · IEEE Trans Neural Netw Learn Syst · 2026

basic_science · Level V

Where this comes from

Abstract

Benefiting from recent advances in vision-language models (VLMs), numerous CLIP-based zero-shot anomaly detection (ZSAD) methods have been proposed to address the cold-start problem. Despite their impressive performance, these methods still depend on manual prompt engineering, and their coarse-grained text prompts struggle to capture the diverse patterns of anomalies, resulting in suboptimal visual-text alignment. To overcome these limitations, we propose VCF-CLIP, a visual context-driven fine-grained prompt learning framework built upon CLIP. The novelties of VCF-CLIP lie in two main aspects. First, we propose the prompt prototype learning (PPL) strategy, which learns a pair of unified prompt prototypes representing general normal and anomalous states in a loss-guided manner, thereby eliminating the need for manual prompt design. Second, we propose a lightweight prompt refinement adapter that dynamically aggregates multiscale and multilevel visual features to iteratively refine the prompt prototypes, enabling the generation of instance-specific prompts enriched with fine-grained information. We conduct extensive experiments on 14 benchmarks across industrial and medical domains, and show that VCF-CLIP outperforms existing state-of-the-art ZSAD methods.