Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on quantitative assessments, such as measuring the size of a tumor or the angle of a joint, from which clinicians draw their own diagnostic conclusions. This quantitative reasoning capability remains underexplored and poorly supported in existing VLMs. In this work, we introduce MedVision, a large-scale dataset and benchmark specifically designed to evaluate and improve VLMs on quantitative medical image analysis. MedVision spans 22 public datasets covering diverse anatomies and modalities, with 29.0K 3D images, 11.2M annotated 2D slices, and 24.3M single-instance annotations. We focus on three representative quantitative tasks: (1) detection of anatomical structures and abnormalities, (2) tumor/lesion (T/L) size estimation, and (3) angle/distance (A/D) measurement. We show that current off-the-shelf VLMs perform poorly on these tasks. However, supervised and reinforcement fine-tuning (RFT) on MedVision significantly enhances performance across detection, T/L size estimation, and A/D measurement, yielding MedVision-V0 as a strong open baseline. In the RFT stage, we design and evaluate the efficacy of process rewards, multiplicative reward composition, and multi-task RFT with curriculum learning. MedVision provides a foundation for developing VLMs with robust quantitative reasoning capabilities in medical imaging.
翻译:暂无翻译