A common assumption in representation learning is that globally well-distributed embeddings support robust and generalizable representations. This focus has shaped both training objectives and evaluation protocols, implicitly treating global geometry as a proxy for representational competence. While global geometry effectively encodes which elements are present, it is often insensitive to how they are composed. We investigate this limitation by testing the ability of geometric metrics to predict compositional binding across a diverse suite of vision encoders. We find that standard geometry-based statistics exhibit near-zero correlation with compositional binding. In contrast, functional sensitivity, as measured by the input--output Jacobian, reliably tracks this capability. We further provide an analytic account showing that this disparity arises from objective design, as existing losses explicitly constrain embedding geometry but leave the local input--output mapping unconstrained. These results suggest that global embedding geometry captures only a partial view of representational competence and establish functional sensitivity as a critical complementary axis for modeling composite structure.
翻译:表征学习中有一个常见假设:全局上分布良好的嵌入能够支撑鲁棒且泛化能力强的表征。这一侧重点既塑造了训练目标,也影响了评估协议,隐含地将全局几何视为表征能力的代理指标。尽管全局几何能有效编码"哪些元素存在",但它往往对"元素如何组合"不敏感。我们通过测试几何度量预测视觉编码器组合绑定能力来探究这一局限性。研究发现,标准几何统计量与组合绑定能力之间的相关系数趋近于零。相比之下,由输入-输出雅可比矩阵度量的功能敏感性能够可靠地追踪该能力。我们进一步从解析角度证明,这种差异源于目标函数设计——现有损失函数显式约束了嵌入几何,却未约束局部输入-输出映射。这些结果表明,全局嵌入几何只能捕捉表征能力的部分视角,而功能敏感性可作为建模复合结构的关键互补维度。