3D human pose estimation from sparse multi-view camera rigs is an essential task for numerous applications, including action recognition, sports analysis, and human-robot interaction. While learned methods dominate the field on benchmarks, they require large annotated datasets; training-free optimization-based methods remain promising as they circumvent 3D supervision by solving a correspondence problem across views from 2D detections. Existing combinatorial formulations rely on pairwise associations to model this correspondence problem and enforce global consistency across views only as a downstream constraint. However, reconciling locally plausible pairwise matches becomes brittle under occlusion and noisy detections, where local errors propagate globally. We propose COMPOSE, which recasts multi-view 3D human pose estimation as a weighted exact-cover optimization over a hypergraph of person hypotheses. Our formulation replaces pairwise association and post-hoc consistency enforcement with a single global combinatorial objective. To address the exponentially large candidate space, we introduce a geometric pruning strategy alongside two complementary solvers: an exact Integer Linear Programming formulation and a scalable relaxation via Belief Propagation. Without any 3D supervision, COMPOSE improves average precision by up to 31 points over the best optimization-based method and 13 points over self-supervised learned methods, demonstrating the effectiveness of higher-order combinatorial association for training-free multi-view 3D human pose estimation.
翻译:从稀疏多视角相机系统中估计三维人体姿态是众多应用中的关键任务,包括动作识别、运动分析及人机交互。尽管基于学习的方法在基准测试中占据主导地位,但它们需要大规模标注数据集;而无需训练的优化方法仍然具有前景,因为它们通过从二维检测结果中求解跨视角对应问题,避免了对三维监督的依赖。现有的组合公式依赖成对关联来建模该对应问题,并将跨视角全局一致性仅作为下游约束加以实施。然而,在遮挡和噪声检测条件下,调和局部合理的成对匹配变得脆弱,局部错误会传播至全局。我们提出COMPOSE,该方法将多视角三维人体姿态估计重新定义为基于人体假设超图的加权精确覆盖优化。我们的公式用单一的全局组合目标替代了成对关联与事后一致性约束。为应对指数级增长的候选空间,我们引入了几何剪枝策略,并配合两种互补求解器:精确整数线性规划公式和通过置信传播实现的可扩展松弛方法。在无需任何三维监督的情况下,COMPOSE相较于最优优化方法平均精度提升最高达31个百分点,相较于自监督学习方法提升13个百分点,证明了高阶组合关联在无需训练的多视角三维人体姿态估计中的有效性。