EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models

NeurIPS, 2026
Hao Wang1,2*, Qiwei Zeng1*, Jinghao Lin3*, Shuchang Ye2, Yuezhe Yang2,4, Yige Peng4, Haoyuan Che1†, Jinman Kim2†, Lei Bi4†
*Equal Contribution †Corresponding author
1Jilin University, 2The University of Sydney, 3Northeastern University, 4Shanghai Jiao Tong University

Abstract

Medical vision-language models can miss subtle lesions because sparse, low-contrast cues are underrepresented when local visual tokens are aggregated. EasyLens is a training-free, plug-and-play representation amplifier for frozen medical VLMs. It constructs a pathology-anatomy prototype space, selects lesion-relevant patches using counterfactual prototype reasoning, and strengthens their representations through morphology-guided residual enhancement. Experiments on multiple medical image datasets and frozen VLM backbones show improved subtle-lesion detection without model fine-tuning.

Highlights

  • Build pathology-anatomy prototypes for comparing suspicious image patches.
  • Select lesion-relevant patches and amplify subtle morphological cues.
  • Improve subtle-lesion recognition without fine-tuning frozen medical VLMs.

Method

Overview of subtle-lesion amplification with EasyLens

EasyLens brings weak subtle-lesion cues into focus in frozen medical VLMs.

EasyBank, EasyTag, and EasyAmplifier workflow

EasyBank construction and inference with EasyTag and EasyAmplifier.

Results

Quantitative Results

ModelReXLIDCAbdomen
Stat.Sel.Gen.Stat.Sel.Gen.Stat.Sel.Gen.
MedGemma1.542.8623.334.4120.4527.8241.9315.1249.0438.18
EasyLens66.6731.115.1530.3036.0945.8652.3355.7740.67

Comparison on subtle-lesion status recognition, region selection, and report generation.

Qualitative Results

EasyLens case study on a small pulmonary nodule

Case study of subtle-lesion perception and lesion-aware report generation.

Citation

@article{wang2026easylens,
  title={EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models},
  author={Wang, Hao and Zeng, Qiwei and Lin, Jinghao and Ye, Shuchang and Yang, Yuezhe and Peng, Yige and Che, Haoyuan and Kim, Jinman and Bi, Lei},
  journal={Advances in Neural Information Processing Systems},
  year={2026},
  url={https://arxiv.org/abs/2606.06379}
}