The design of communication systems dedicated to machine learning tasks is one key aspect of goal-oriented communications. In this framework, this article investigates the interplay between data reconstruction and learning from the same compressed observations, particularly focusing on the regression problem. We establish achievable rate-generalization error regions for both parametric and non-parametric regression, where the generalization error measures the regression performance on previously unseen data. The analysis covers both asymptotic and finite block-length regimes, providing fundamental results and practical insights for the design of coding schemes dedicated to regression. The asymptotic analysis relies on conventional Wyner-Ziv coding schemes which we extend to study the convergence of the generalization error. The finite-length analysis uses the notions of information density and dispersion with additional term for the generalization error. We further investigate the trade-off between reconstruction and regression in both asymptotic and non-asymptotic regimes. Contrary to the existing literature which focused on other learning tasks, our results state that in the case of regression, there is no trade-off between data reconstruction and regression in the asymptotic regime. We also observe the same absence of trade-off for the considered achievable scheme in the finite-length regime, by analyzing correlation between distortion and generalization error.
翻译:面向机器学习任务的通信系统设计是指向性通信的关键方面之一。在此框架下,本文研究了从同一压缩观测数据中进行数据重构与学习之间的交互关系,特别聚焦于回归问题。我们为参数化和非参数化回归建立了可达速率-泛化误差区域,其中泛化误差衡量回归在未见过数据上的性能。分析涵盖了渐近和有限块长两种场景,为专用于回归的编码方案设计提供了基础性结论和实践见解。渐近分析基于经典Wyner-Ziv编码方案,我们将其拓展以研究泛化误差的收敛性。有限长分析运用了信息密度和色散的概念,并引入泛化误差的附加项。我们进一步研究了渐近与非渐近机制下重构与回归之间的权衡。与聚焦于其他学习任务的现有文献相反,我们的结果表明:在回归情形下,渐近机制中数据重构与回归之间不存在权衡。通过分析失真与泛化误差之间的相关性,我们观察到在有限长机制下所考虑的可达方案同样不存在这种权衡。