Publikasjonsdetaljer
- Journal: Lecture Notes in Computer Science, vol. LNCS 16879, 2026
-
Lenke:
- ARKIV: hdl.handle.net/11250/5580926
Vision–language models (VLMs) are increasingly used in medical imaging, yet their robustness to spurious correlations remains insufficiently characterized. We propose a controlled evaluation framework that uses synthetic artifacts modelled after common acquisition confounders to parametrically vary correlation strength, and pairs two complementary test protocols — artifact removal and artifact inversion — to isolate whether models rely on clinical features or visual shortcuts. Applying the framework to diabetic retinopathy grading in fundus photography and BI-RADS-based assessment in mammography, we evaluate five architectures spanning a spectrum from no concept supervision to full multi-level image–concept alignment. We find that VLM backbones retain clinical signal when shortcuts are absent, yet actively follow spurious associations when they conflict with pathology, degrading faster than standard visual backbones — a dual encoding that is only exposed when evaluation goes beyond clean test sets. Among concept-based strategies, only architectures that both reshape the feature space toward clinical concepts and shield the classifier from non-clinical signal provide meaningful resilience. The framework is architecture-agnostic and applicable to any vision or multimodal model. To support evaluation in mammography, where the combinatorial richness of the BI-RADS lexicon cannot be feasibly captured by binary concepts alone, we release expert annotations with pixel-wise delineation of findings for 400 images. They can be found, alongside the code, at the following href{https://github.com/valecorbetta/framework_for_vlm_eval}{link}.