Medical vision-language models can miss subtle lesions because sparse, low-contrast cues are underrepresented when local visual tokens are aggregated. EasyLens is a training-free, plug-and-play representation amplifier for frozen medical VLMs. It constructs a pathology-anatomy prototype space, selects lesion-relevant patches using counterfactual prototype reasoning, and strengthens their representations through morphology-guided residual enhancement. Experiments on multiple medical image datasets and frozen VLM backbones show improved subtle-lesion detection without model fine-tuning.
EasyLens brings weak subtle-lesion cues into focus in frozen medical VLMs.
EasyBank construction and inference with EasyTag and EasyAmplifier.
| Model | ReX | LIDC | Abdomen | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Stat. | Sel. | Gen. | Stat. | Sel. | Gen. | Stat. | Sel. | Gen. | |
| MedGemma1.5 | 42.86 | 23.33 | 4.41 | 20.45 | 27.82 | 41.93 | 15.12 | 49.04 | 38.18 |
| EasyLens | 66.67 | 31.11 | 5.15 | 30.30 | 36.09 | 45.86 | 52.33 | 55.77 | 40.67 |
Comparison on subtle-lesion status recognition, region selection, and report generation.
Case study of subtle-lesion perception and lesion-aware report generation.
@article{wang2026easylens,
title={EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models},
author={Wang, Hao and Zeng, Qiwei and Lin, Jinghao and Ye, Shuchang and Yang, Yuezhe and Peng, Yige and Che, Haoyuan and Kim, Jinman and Bi, Lei},
journal={Advances in Neural Information Processing Systems},
year={2026},
url={https://arxiv.org/abs/2606.06379}
}