Part-aware 3D generation aims to synthesize structured objects with semantically meaningful components, yet often suffers from structural ambiguity due to identity-layout entanglement. Existing methods either infer part identity and spatial layout implicitly, which can lead to unstable part allocation (e.g., slot swapping or part merging), or rely on strong layout conditions that are difficult to obtain in practice. We attribute this ambiguity to identity-slot permutation freedom: without explicit identity-slot alignment, the correspondence between semantic parts and generation slots is not identifiable during training, allowing multiple slot assignments to fit the same supervision and leading to inconsistent decomposition. Based on this insight, we argue that stable part-aware generation requires identity-aligned one-to-one slot modelling. We therefore propose an identity-slot aligned framework, ISAP-3D, which anchors each part with semantic identity tokens and performs identity-conditioned one-to-one layout prediction, followed by layout-conditioned geometry synthesis. Structured local-global conditioning maintains identity alignment across semantic, spatial, and geometric stages. We also construct a part-level dataset with a unified semantic protocol to enable learnable and consistent identity-slot alignment. Extensive experiments demonstrate improved structural stability, controllability, and robustness over state-of-the-art part-aware generation baselines.
翻译:部件感知三维生成旨在合成具有语义意义组件的结构化物体,但常因身份与布局的纠缠导致结构歧义。现有方法要么隐式推断部件身份与空间布局,易引发不稳定的部件分配(如槽交换或部件合并),要么依赖实践中难以获取的强布局条件。我们将此歧义归因于身份-槽排列自由度:缺乏显式的身份-槽对齐,训练过程中语义部件与生成槽之间的对应关系无法辨识,允许多种槽分配适配同一监督信号,导致分解结果不一致。基于这一见解,我们认为稳定的部件感知生成需要身份对齐的一对一槽建模。为此,我们提出身份-槽对齐框架ISAP-3D,为每个部件锚定语义身份标记,执行身份条件的一对一布局预测,随后进行布局条件的几何合成。结构化的局部-全局条件机制在语义、空间和几何阶段保持身份对齐。我们同时构建了部件级数据集并采用统一语义协议,以实现可学习的、一致的身份-槽对齐。大量实验表明,与最先进的部件感知生成基线相比,该方法在结构稳定性、可控性和鲁棒性方面均取得显著提升。