Multi-task learning is frequently used to model a set of related response variables from the same set of features, improving predictive performance and modeling accuracy relative to methods that handle each response variable separately. Despite the potential of multi-task learning to yield more powerful inference than single-task alternatives, prior work in this area has largely omitted uncertainty quantification. Our focus in this paper is a common multi-task problem in neuroimaging, where the goal is to understand the relationship between multiple cognitive task scores (or other subject-level assessments) and brain connectome data collected from imaging. We propose a framework for selective inference to address this problem, with the flexibility to: (i) jointly identify the relevant covariates for each task through a sparsity-inducing penalty, and (ii) conduct valid inference in a model based on the estimated sparsity structure. Our framework offers a new conditional procedure for inference, based on a refinement of the selection event that yields a tractable selection-adjusted likelihood. This gives an approximate system of estimating equations for maximum likelihood inference, solvable via a single convex optimization problem, and enables us to efficiently form confidence intervals with approximately the correct coverage. Applied to both simulated data and data from the Adolescent Brain Cognitive Development (ABCD) study, our selective inference methods yield tighter confidence intervals than commonly used alternatives, such as data splitting. We also demonstrate through simulations that multi-task learning with selective inference can more accurately recover true signals than single-task methods.
翻译:多任务学习常用于对同一组特征相关的多个响应变量进行建模,相较于单独处理每个响应变量的方法,能够提升预测性能和建模精度。尽管多任务学习相比单任务方法具有产生更强大推断的潜力,但先前该领域的研究在很大程度上忽略了不确定性量化。本文聚焦于神经影像学中一个常见的多任务问题,其目标是理解多个认知任务得分(或其他受试者评估指标)与影像收集的脑连接组数据之间的关系。我们提出了一种用于解决此问题的选择性推断框架,具有以下灵活性:(i) 通过稀疏性惩罚联合识别每个任务的相关协变量,以及(ii) 在基于估计稀疏结构的模型中进行有效推断。该框架提供了一种新的基于条件推理的推断程序,通过对选择事件进行精炼,得到可处理的经选择调整的似然函数。这给出了一套近似估计方程系统用于最大似然推断,可通过单个凸优化问题进行求解,并使我们能够高效地构建具有近似正确覆盖率的置信区间。将我们的选择性推断方法应用于模拟数据以及青少年脑认知发展(ABCD)研究的数据,结果表明,与数据分裂等常用替代方法相比,该方法可产生更紧凑的置信区间。我们还通过模拟实验证明,采用选择性推断的多任务学习比单任务方法能更准确地恢复真实信号。