In bimanual robotic manipulation, task-relevant visual information varies with the task stage and context, while the interaction of the two arms shifts between independent and coordinated modes, making policy learning challenging. However, existing monolithic Vision-Language-Action (VLA) policies process diverse visual inputs and interaction patterns through a single shared representation and action generation pathway, often failing to separately account for visual relevance and bimanual interaction structure. To address this issue, we propose a bimanual manipulation VLA framework based on Dual-Level Structural Decomposition. The View-Selective Visual Router dynamically adjusts wrist-view contributions to emphasize relevant visual cues, while the Interaction-Aware Action Mixture-of-Experts (MoE) decomposes action generation into coordinated and arm-wise pathways to adapt to varying bimanual interaction modes. We evaluate the proposed method on six simulated bimanual manipulation tasks in RoboTwin 2.0 and three long-horizon real-world tasks. Our model improves the overall average success rate over a monolithic baseline by 27.7% in simulation and 43.3% in real-world evaluation, while consistently outperforming single-module variants across both settings. These results demonstrate that jointly considering selective visual processing and explicit decomposition of bimanual interaction structures provides an effective inductive bias for robust bimanual manipulation.
翻译:在双臂机器人操作中,任务相关的视觉信息随任务阶段和上下文而变化,同时双臂的交互在独立与协同模式之间切换,这使得策略学习具有挑战性。然而,现有的单体型视觉-语言-动作(VLA)策略通过单一共享表示和动作生成路径处理多样化的视觉输入与交互模式,通常无法分别考虑视觉相关性及双臂交互结构。为解决这一问题,我们提出了一种基于双层级结构分解的双臂操作VLA框架。其中,视角选择性视觉路由器动态调整腕部视角的贡献以强调相关视觉线索,而交互感知动作混合专家(MoE)将动作生成分解为协同与独立手臂通路,以适配不同的双臂交互模式。我们在RoboTwin 2.0中的六个模拟双臂操作任务和三个长时域真实世界任务上评估了所提方法。相比单体型基线,我们的模型在模拟环境中将整体平均成功率提升了27.7%,在真实世界评估中提升了43.3%,且在这两种设置下均一致优于单模块变体。这些结果表明,联合考虑选择性视觉处理与双臂交互结构的显式分解,为鲁棒的双臂操作提供了有效的归纳偏置。