Reconstructing articulated 3D objects is important for animation, gaming, and robotic simulations. Recent neural networks can estimate the articulated structure of 3D objects, but their generalization remains limited by the scarcity of annotated data for this task. To address this gap, we introduce Instruct-Particulate, a model that takes a 3D mesh together with a target kinematic specification, including part descriptions, connectivity, joint types, and optional point prompts, and predicts the corresponding kinematic part segmentation and joint motion parameters. The kinematic specification disambiguates the task and allows the model to target annotations of different granularity, thereby making it possible to use more abundant heterogeneous training data. At test time, the kinematic specification can be obtained automatically from large-scale vision-language models, so the model can be applied to any input mesh. To train our model at scale, we construct a heterogeneous dataset of more than 150,000 articulated 3D objects, extending existing publicly available collections with data obtained by partially labelling other 3D models (monolithic or already decomposed into parts) with kinematic labels by means of vision-language models. Experiments show that our model generalizes better across categories and to AI-generated meshes, enabling articulated asset reconstruction from real-world images via image-to-3D models.
翻译:重构可关节化3D物体对于动画、游戏与机器人仿真具有重要意义。近期神经网络虽能估计3D物体的关节结构,但其泛化能力受限于该任务标注数据的稀缺性。为弥补这一不足,我们提出指令-颗粒(Instruct-Particulate)模型:该模型以3D网格及目标运动学规格(含部件描述、连接关系、关节类型及可选点提示)为输入,预测对应的运动学部件分割与关节运动参数。运动学规格消除了任务歧义性,使模型可针对不同粒度的标注目标,从而利用更丰富的异构训练数据。测试阶段,运动学规格可通过大规模视觉-语言模型自动获取,因此该模型可应用于任意输入网格。为规模化训练模型,我们构建了包含超过15万个可关节化3D物体的异构数据集:通过视觉-语言模型对已有公开数据及部分标注其他3D模型(单体或已分解为部件的模型)的运动学标签进行扩充。实验表明,相比现有方法,本模型在跨类别及AI生成网格中具有更优的泛化性能,能借助图像到3D模型实现真实世界图像的可关节化资产重建。