Industrial products such as valves and circuit breakers are defined by dense technical specifications that govern procurement, compatibility, and safety across supply chains. These specifications are scattered across multiple heterogeneous product images, including specification tables, nameplates, and technical drawings, yet whether Multimodal Large Language Models (MLLMs) can reliably recover them remains underexplored. To fill this gap, we introduce IndustryBench-MIPU, the first large-scale benchmark for multi-image industrial product understanding, built around structured attribute extraction -- recovering property-value pairs from product images. This task jointly probes text recognition on specification tables and nameplates, visual reasoning over technical drawings, domain knowledge to decode industrial terminology, and cross-image evidence integration to assemble scattered specifications. Concretely, the benchmark comprises 4,559 products across 27,652 images with 103,703 annotations spanning 18 industrial categories, constructed through multi-model consensus and three-tier quality assurance. Evaluating nine MLLMs under both single-image and product-level multi-image settings reveals a stark completeness gap: models achieve high precision (86--94%) but the best recovers only 49.9% of product-level attributes; moving from single-image to multi-image extraction costs 15--34 percentage points of recall. Multi-image completeness, not single-image accuracy, is the core bottleneck. Dataset and code are publicly available.


翻译:阀门、断路器等工业产品由密集的技术规格定义,这些规格支配着供应链中的采购、兼容性和安全性。这些技术规格分散在多个异构的产品图像中,包括规格表、铭牌和技术图纸,然而,多模态大语言模型(MLLMs)能否可靠地恢复它们仍未得到充分探索。为填补这一空白,我们提出了IndustryBench-MIPU,这是首个用于多图像工业产品理解的大规模基准,其核心是结构化属性提取——从产品图像中恢复属性-值对。该任务联合探究了规格表和铭牌上的文本识别、技术图纸上的视觉推理、解码工业术语的领域知识,以及跨图像证据集成以整合分散的规格。具体而言,该基准包含来自27,652张图像的4,559个产品,涵盖103,703条标注,跨越18个工业类别,通过多模型共识和三层质量保证构建。在单图像和产品级多图像设置下评估九种MLLMs,揭示了一个显著的不完整性差距:模型实现了高精确度(86-94%),但最佳模型仅恢复了49.9%的产品级属性;从单图像到多图像提取,召回率下降了15-34个百分点。多图像完整性(而非单图像准确性)是核心瓶颈。数据集和代码已公开提供。

0
下载
关闭预览

相关内容

用来满足人们需求和欲望的物体或无形的载体。好的产品大家都喜欢
MME-Survey:多模态大型语言模型评估的综合性调查
专知会员服务
43+阅读 · 2024年12月1日
多模态大规模语言模型基准的综述
专知会员服务
41+阅读 · 2024年8月25日
使用多模态语言模型生成图像
专知会员服务
32+阅读 · 2023年8月23日
【CVPR2023】基于混合融合的多模态工业异常检测
专知会员服务
46+阅读 · 2023年3月6日
【PHM】NIST:PHM制造工艺流程技术和指标路线图
产业智能官
11+阅读 · 2019年1月13日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
6+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员