Single-Index Models are high-dimensional regression problems with planted structure, whereby labels depend on an unknown one-dimensional projection of the input via a generic, non-linear, and potentially non-deterministic transformation. As such, they encompass a broad class of statistical inference tasks, and provide a rich template to study statistical and computational trade-offs in the high-dimensional regime. While the information-theoretic sample complexity to recover the hidden direction is linear in the dimension $d$, we show that computationally efficient algorithms, both within the Statistical Query (SQ) and the Low-Degree Polynomial (LDP) framework, necessarily require $\Omega(d^{k^\star/2})$ samples, where $k^\star$ is a "generative" exponent associated with the model that we explicitly characterize. Moreover, we show that this sample complexity is also sufficient, by establishing matching upper bounds using a partial-trace algorithm. Therefore, our results provide evidence of a sharp computational-to-statistical gap (under both the SQ and LDP class) whenever $k^\star>2$. To complete the study, we provide examples of smooth and Lipschitz deterministic target functions with arbitrarily large generative exponents $k^\star$.
翻译:单指标模型是一类具有植入结构的高维回归问题,其标签依赖于输入未知一维投影经通用非线性且可能非确定性变换后的结果。因此,这类模型涵盖了广泛的统计推断任务,并为研究高维场景下的统计-计算权衡提供了丰富模板。尽管恢复隐藏方向的信息论样本复杂度与维度$d$呈线性关系,但我们证明:在统计查询(SQ)和低阶多项式(LDP)框架下,计算高效算法必然需要$\Omega(d^{k^\star/2})$个样本,其中$k^\star$是我们明确表征的模型"生成性"指数。进一步,我们通过部分迹算法建立匹配上界,证明了该样本复杂度亦具有充分性。因此,当$k^\star>2$时,我们的结果揭示了尖锐的计算-统计鸿沟(在SQ与LDP类下均成立)。为完善研究,我们给出了具有任意大生成指数$k^\star$的Lipschitz光滑确定性目标函数实例。