The release of SAM 3D Body is a recent development in human mesh recovery, demonstrating improved performance in producing clean, topologically coherent meshes from single images. By leveraging the Momentum Human Rig (MHR), it achieves robustness to occlusion and diverse poses. However, our evaluation reveals a specific and consistent limitation: the model struggles to reconstruct detailed anthropometric deviations, particularly in populations exhibiting distinctive morphological alterations such as geriatric muscle atrophy, scoliosis, or pregnancy, even when these features are prominent in the input image. In this paper, we investigate this phenomenon not as a failure of the model's capacity, but as a byproduct of the "perception-distortion trade-off". We posit that the architectural reliance on the low-dimensional parametric MHR representation, combined with semantic-invariant conditioning (DINOv3) and annotation-based alignment, creates a pervasive "regression to the mean" effect. We analyze these mechanisms to understand why individual biological details are smoothed out. Furthermore, we state our contributions by proposing specific, constructive pathways for future work, such as implicit-explicit hybrid representations and Medical-in-the-Loop alignment, to extend the baseline performance of SAM 3D Body into the high-precision medical domain.
翻译:SAM 3D人体模型的发布是人体网格重建领域的最新进展,其在从单张图像生成干净、拓扑一致的人体网格方面展现了更优性能。通过利用动量人体骨架(MHR),该模型实现了对遮挡和多样化姿态的鲁棒性。然而,我们的评估揭示了一个特定且一致的局限性:该模型难以重建详细的人体测量偏差,尤其是在表现出显著形态学改变的人群中(例如老年性肌肉萎缩、脊柱侧弯或妊娠状态),即使这些特征在输入图像中十分突出。本文研究这一现象,并非将其视为模型能力的不足,而是"感知-失真权衡"的副产品。我们认为,模型架构对低维参数化MHR表示的依赖,结合语义不变条件(DINOv3)与基于标注的对齐,产生了普遍存在的"回归均值"效应。我们分析了这些机制,以理解个体生物细节被平滑化的原因。此外,我们通过提出具体的建设性未来研究方向(如隐式-显式混合表示和医学在环对齐)阐明贡献,旨在将SAM 3D人体模型的基线性能拓展至高精度医学领域。