We consider the problem of linear regression with self-selection bias in the unknown-index setting, as introduced in recent work by Cherapanamjeri, Daskalakis, Ilyas, and Zampetakis [STOC 2023]. In this model, one observes $m$ i.i.d. samples $(\mathbf{x}_{\ell},z_{\ell})_{\ell=1}^m$ where $z_{\ell}=\max_{i\in [k]}\{\mathbf{x}_{\ell}^T\mathbf{w}_i+\eta_{i,\ell}\}$, but the maximizing index $i_{\ell}$ is unobserved. Here, the $\mathbf{x}_{\ell}$ are assumed to be $\mathcal{N}(0,I_n)$ and the noise distribution $\mathbf{\eta}_{\ell}\sim \mathcal{D}$ is centered and independent of $\mathbf{x}_{\ell}$. We provide a novel and near optimally sample-efficient (in terms of $k$) algorithm to recover $\mathbf{w}_1,\ldots,\mathbf{w}_k\in \mathbb{R}^n$ up to additive $\ell_2$-error $\varepsilon$ with polynomial sample complexity $\tilde{O}(n)\cdot \mathsf{poly}(k,1/\varepsilon)$ and significantly improved time complexity $\mathsf{poly}(n,k,1/\varepsilon)+O(\log(k)/\varepsilon)^{O(k)}$. When $k=O(1)$, our algorithm runs in $\mathsf{poly}(n,1/\varepsilon)$ time, generalizing the polynomial guarantee of an explicit moment matching algorithm of Cherapanamjeri, et al. for $k=2$ and when it is known that $\mathcal{D}=\mathcal{N}(0,I_k)$. Our algorithm succeeds under significantly relaxed noise assumptions, and therefore also succeeds in the related setting of max-linear regression where the added noise is taken outside the maximum. For this problem, our algorithm is efficient in a much larger range of $k$ than the state-of-the-art due to Ghosh, Pananjady, Guntuboyina, and Ramchandran [IEEE Trans. Inf. Theory 2022] for not too small $\varepsilon$, and leads to improved algorithms for any $\varepsilon$ by providing a warm start for existing local convergence methods.
翻译:我们考虑在未知索引设定下具有自选择偏差的线性回归问题,该问题由Cherapanamjeri、Daskalakis、Ilyas和Zampetakis在近期工作中引入(STOC 2023)。在该模型中,观测到$m$个独立同分布样本$(\mathbf{x}_{\ell},z_{\ell})_{\ell=1}^m$,其中$z_{\ell}=\max_{i\in [k]}\{\mathbf{x}_{\ell}^T\mathbf{w}_i+\eta_{i,\ell}\}$,但最大化索引$i_{\ell}$未被观测。这里假设$\mathbf{x}_{\ell}\sim \mathcal{N}(0,I_n)$,且噪声分布$\mathbf{\eta}_{\ell}\sim \mathcal{D}$为中心化的,并与$\mathbf{x}_{\ell}$独立。我们提出了一种新颖且近似最优样本高效(关于$k$)的算法,能以多项式样本复杂度$\tilde{O}(n)\cdot \mathsf{poly}(k,1/\varepsilon)$恢复$\mathbf{w}_1,\ldots,\mathbf{w}_k\in \mathbb{R}^n$至加性$\ell_2$误差$\varepsilon$,并显著提升时间复杂度至$\mathsf{poly}(n,k,1/\varepsilon)+O(\log(k)/\varepsilon)^{O(k)}$。当$k=O(1)$时,算法运行时间为$\mathsf{poly}(n,1/\varepsilon)$,推广了Cherapanamjeri等人针对$k=2$且已知$\mathcal{D}=\mathcal{N}(0,I_k)$情形下显式矩匹配算法的多项式保证。我们的算法在显著放宽的噪声假设下仍能成功,因此也适用于加性噪声置于最大值外部的最大线性回归相关设定。针对该问题,对于非极小$\varepsilon$情形,我们的算法在比Ghosh、Pananjady、Guntuboyina和Ramchandran(IEEE Trans. Inf. Theory 2022)提出的现有技术更广泛的$k$范围内保持高效,并通过为现有局部收敛方法提供热启动,改进了任意$\varepsilon$下的算法性能。