We present `GL-LowPopArt`, a novel Catoni-style estimator for generalized low-rank trace regression. Building on `LowPopArt` (Jang et al., 2024), it employs a two-stage approach: nuclear norm regularization followed by matrix Catoni estimation. We establish state-of-the-art estimation error bounds, surpassing existing guarantees (Fan et al., 2019; Kang et al., 2022), and reveal a novel experimental design objective, $\mathrm{GL}(π)$. The key technical challenge is controlling bias from the nonlinear inverse link function, which we address with our two-stage approach. We prove a *local minimax lower bound*, showing that our `GL-LowPopArt` enjoys instance-wise optimality up to the condition number of the ground-truth Hessian. Our method immediately achieves an improved Frobenius error guarantee for generalized linear matrix completion. We also introduce a new problem setting called **bilinear dueling bandits**, a contextualized version of dueling bandits with a general preference model. Using an explore-then-commit approach with `GL-LowPopArt', we show an improved Borda regret bound over naïve vectorization (Wu et al., 2024).
翻译:我们提出`GL-LowPopArt`,一种面向广义低秩迹回归的新型Catoni型估计量。该估计量基于`LowPopArt`(Jang等,2024)构建,采用两阶段方法:核范数正则化后接矩阵Catoni估计。我们建立了超越现有理论保证(Fan等,2019;Kang等,2022)的最优估计误差界,并揭示了一个新颖的实验设计目标函数$\mathrm{GL}(π)$。关键技术挑战在于控制非线性逆链接函数引起的偏差,我们通过两阶段方法解决了该问题。我们证明了*局部极小极大下界*,表明所提出的`GL-LowPopArt`在真实Hessian矩阵条件数范围内具有实例级最优性。该方法可直接改进广义线性矩阵补全的Frobenius误差保证。我们还引入了名为**双线性对决多臂赌博机**的新问题设定——一种具有通用偏好模型的情境化对决赌博机。通过采用`GL-LowPopArt`的探索-然后-承诺方法,我们证明相较于朴素向量化方法(Wu等,2024)具有更优的Borda遗憾界。