We consider information-theoretic bounds on expected generalization error for statistical learning problems in a networked setting. In this setting, there are $K$ nodes, each with its own independent dataset, and the models from each node have to be aggregated into a final centralized model. We consider both simple averaging of the models as well as more complicated multi-round algorithms. We give upper bounds on the expected generalization error for a variety of problems, such as those with Bregman divergence or Lipschitz continuous losses, that demonstrate an improved dependence of $1/K$ on the number of nodes. These "per node" bounds are in terms of the mutual information between the training dataset and the trained weights at each node, and are therefore useful in describing the generalization properties inherent to having communication or privacy constraints at each node.
翻译:我们考虑网络环境下统计学习问题中期望泛化误差的信息论界。在该设置中,存在 $K$ 个节点,每个节点拥有独立数据集,且各节点模型需聚合为最终的中心化模型。我们同时考虑简单的模型平均方法以及更复杂的多轮算法。针对多种问题(例如具有Bregman散度或Lipschitz连续损失函数的问题),我们给出了期望泛化误差的上界,这些上界在节点数量上表现出改进的 $1/K$ 依赖关系。这些"每节点"界以各节点训练数据集与训练权重之间的互信息来表述,因此有助于刻画因节点通信或隐私约束所固有的泛化特性。