X-ray contraband detection is critical for security in large-scale logistics and transportation, yet conventional detectors struggle to adapt to emerging contraband types and lack fundamental visual understanding. Vision-language models (VLMs) offer strong generalization but are hindered by the scarcity of high-quality X-ray image-caption data. To bridge this critical gap, we present MMXray, a meticulously curated benchmark of 52,124 image-caption pairs spanning 28 fine-grained classes of X-ray contraband. To enrich MMXray with realistic occlusion patterns, we further introduce CleanDET, a dedicated synthesis dataset containing clean foreground contraband images from 28 categories and background images with diverse density levels, together with AnyContraSyn, a controllable synthesis method designed to operate on CleanDET. We also develop OnePipe, an extensible pipeline for systematic data curation. Built on MMXray, we propose OneFocus, a unified VLM that supports four core tasks: visual question answering, contraband localization, classification, and image understanding. OneFocus achieves state-of-the-art performance in X-ray contraband understanding and demonstrates robust cross-domain generalization, establishing a strong vision-language baseline for security screening.
翻译:X光违禁品检测对大规模物流与运输安全至关重要,但传统检测器难以适应新型违禁品类别且缺乏基础视觉理解能力。视觉-语言模型虽具备强大泛化能力,却受限于高质量X光图像-文本配对数据的稀缺性。为填补这一关键空白,我们提出了MMXray——一个精心构建的基准数据集,包含52,124组图像-文本对,涵盖28个细粒度X光违禁品类别。为模拟真实遮挡模式,我们进一步引入CleanDET专用合成数据集(包含28类纯净前景违禁品图像及多密度背景图像),并设计了可控合成方法AnyContraSyn。同时开发了OnePipe可扩展数据系统化处理流水线。基于MMXray数据集,我们提出OneFocus统一视觉-语言模型,可支持四项核心任务:视觉问答、违禁品定位、分类与图像理解。OneFocus在X光违禁品理解任务中达到最优性能,并展现出稳健的跨域泛化能力,为安检场景建立了强大的视觉-语言基线。