This paper studies the Minimal Embeddable Dimension (MED): the least dimension in which there exists a configuration of $m$ object vectors so that every subset of size at most $k$ is exactly retrieved by score comparison. Our result shows MED is $Θ(k)$, independent of $m$, for inner product, Euclidean distance, and cosine similarity. We then consider Robust MED (RMED), where all vectors are unit normed and an $ε$ gap of scores is required. We derive the $m$-dependent feasibility ceiling $ε_\star(m,k)=m/\sqrt{k(m-1)(m-k)}$, which approaches $1/\sqrt{k}$ when $m\gg k$, and a Gaussian centroid construction gives a robust witness upper bound in the feasible margin regime. Numerical simulation on synthetic top-$2$ retrieval with cyclic polytope and centroid query optimization confirmed our theoretical claims. Experiments on LIMIT and LIMIT-small datasets also show that simple embedding-based retrieval baselines can overfit and outperform the reported single-vector LLM embedding baseline. Both theoretical and empirical findings rule out the lack of exact geometric capacity as the obstruction.
翻译:本文研究了最小可嵌入维度(MED):即存在\(m\)个对象向量的配置,使得每个大小至多为\(k\)的子集都能通过分数比较被精确检索到的最小维度。我们的结果表明,对于内积、欧氏距离和余弦相似度,MED为\(\Theta(k)\),且与\(m\)无关。随后我们考虑了鲁棒最小可嵌入维度(RMED),其中所有向量均为单位范数,且要求分数存在\(\epsilon\)间隔。我们推导出与\(m\)相关的可行性上限\(\epsilon_\star(m,k)=m/\sqrt{k(m-1)(m-k)}\),当\(m\gg k\)时该值趋近于\(1/\sqrt{k}\),同时在高斯质心构造下给出了可行余量范围内的鲁棒性上界。基于循环多面体和质心查询优化的合成Top-2检索数值实验验证了我们的理论结论。在LIMIT和LIMIT-small数据集上的实验还表明,简单的基于嵌入的检索基线可能过拟合,并优于已报道的单向量LLM嵌入基线。理论分析与实证结果共同排除了精确几何容量不足作为障碍的可能性。