We present new results to model and understand the role of encoder-decoder design in machine learning (ML) from an information-theoretic angle. We use two main information concepts, information sufficiency (IS) and mutual information loss (MIL), to represent predictive structures in machine learning. Our first main result provides a functional expression that characterizes the class of probabilistic models consistent with an IS encoder-decoder latent predictive structure. This result formally justifies the encoder-decoder forward stages many modern ML architectures adopt to learn latent (compressed) representations for classification. To illustrate IS as a realistic and relevant model assumption, we revisit some known ML concepts and present some interesting new examples: invariant, robust, sparse, and digital models. Furthermore, our IS characterization allows us to tackle the fundamental question of how much performance (predictive expressiveness) could be lost, using the cross entropy risk, when a given encoder-decoder architecture is adopted in a learning setting. Here, our second main result shows that a mutual information loss quantifies the lack of expressiveness attributed to the choice of a (biased) encoder-decoder ML design. Finally, we address the problem of universal cross-entropy learning with an encoder-decoder design where necessary and sufficiency conditions are established to meet this requirement. In all these results, Shannon's information measures offer new interpretations and explanations for representation learning.
翻译:本文从信息论角度出发,为建模和理解机器学习中编码器-解码器设计的作用提出了新的研究成果。我们采用两个核心信息概念——信息充分性(IS)与互信息损失(MIL)——来表征机器学习中的预测结构。第一个主要成果给出了一个函数表达式,用于刻画符合IS编码器-解码器潜在预测结构的概率模型类别。该结果从形式上证明了现代许多机器学习架构为学习分类所需的潜在(压缩)表示而采用的编码器-解码器前向阶段的合理性。为说明IS作为一种现实且相关的模型假设,我们重新审视了一些已知的机器学习概念,并提出了若干有趣的新示例:不变模型、鲁棒模型、稀疏模型与数字模型。此外,我们的IS表征使我们能够探究一个基础问题:在学习场景中采用特定编码器-解码器架构时,以交叉熵风险衡量的性能(预测表达能力)可能损失多少。在此,我们的第二个主要结果表明,互信息损失可量化因选择(有偏的)编码器-解码器机器学习设计所导致的表达能力缺失。最后,我们探讨了采用编码器-解码器设计的通用交叉熵学习问题,并建立了满足该要求的必要与充分条件。在所有结果中,香农信息测度为表征学习提供了新的解释与阐释。