站内搜索 - AI4Meder

论文ICLR 2026 Poster2026 年clinical NLP

VLM-SubtleBench：VLM 距离人类级细微比较推理还有多远？

ICLR 2026 Poster accepted paper at ICLR 2026. The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for vision-language models (VLMs) have recently emerged, they primarily focus on images with large, salient differences and fail to capture the nuanced reasoning required for real-world applications. In this work, we introduce **VLM-SubtleBench**, a benchmark designed to evaluate VLMs on *subtle comparative reasoning*. Our benchmark covers ten difference types—Attribute, State, Emotion, Temporal, Spatial, Existence, Quantity, Quality, Viewpoint, and Action—and curate paired question–image sets reflecting these fine-grained variations.

医学影像计算医疗多模态临床语言智能论文 Vision-language Models Multimodal Large Language Models 查看论文详情

搜索医学 AI 论文与资源

1 条结果

VLM-SubtleBench：VLM 距离人类级细微比较推理还有多远？