Industrial products such as valves and circuit breakers are defined by dense technical specifications that govern procurement, compatibility, and safety across supply chains. These specifications are scattered across multiple heterogeneous product images, including specification tables, nameplates, and technical drawings, yet whether Multimodal Large Language Models (MLLMs) can reliably recover them remains underexplored. To fill this gap, we introduce IndustryBench-MIPU, the first large-scale benchmark for multi-image industrial product understanding, built around structured attribute extraction -- recovering property-value pairs from product images. This task jointly probes text recognition on specification tables and nameplates, visual reasoning over technical drawings, domain knowledge to decode industrial terminology, and cross-image evidence integration to assemble scattered specifications. Concretely, the benchmark comprises 4,559 products across 27,652 images with 103,703 annotations spanning 18 industrial categories, constructed through multi-model consensus and three-tier quality assurance. Evaluating nine MLLMs under both single-image and product-level multi-image settings reveals a stark completeness gap: models achieve high precision (86--94%) but the best recovers only 49.9% of product-level attributes; moving from single-image to multi-image extraction costs 15--34 percentage points of recall. Multi-image completeness, not single-image accuracy, is the core bottleneck. Dataset and code are publicly available.
翻译:阀门、断路器等工业产品由密集的技术规格定义,这些规格支配着供应链中的采购、兼容性和安全性。这些技术规格分散在多个异构的产品图像中,包括规格表、铭牌和技术图纸,然而,多模态大语言模型(MLLMs)能否可靠地恢复它们仍未得到充分探索。为填补这一空白,我们提出了IndustryBench-MIPU,这是首个用于多图像工业产品理解的大规模基准,其核心是结构化属性提取——从产品图像中恢复属性-值对。该任务联合探究了规格表和铭牌上的文本识别、技术图纸上的视觉推理、解码工业术语的领域知识,以及跨图像证据集成以整合分散的规格。具体而言,该基准包含来自27,652张图像的4,559个产品,涵盖103,703条标注,跨越18个工业类别,通过多模型共识和三层质量保证构建。在单图像和产品级多图像设置下评估九种MLLMs,揭示了一个显著的不完整性差距:模型实现了高精确度(86-94%),但最佳模型仅恢复了49.9%的产品级属性;从单图像到多图像提取,召回率下降了15-34个百分点。多图像完整性(而非单图像准确性)是核心瓶颈。数据集和代码已公开提供。