Constructing useful representations across a large number of tasks is a key requirement for sample-efficient intelligent systems. A traditional idea in multitask learning (MTL) is building a shared representation across tasks which can then be adapted to new tasks by tuning last layers. A desirable refinement of using a shared one-fits-all representation is to construct task-specific representations. To this end, recent PathNet/muNet architectures represent individual tasks as pathways within a larger supernet. The subnetworks induced by pathways can be viewed as task-specific representations that are composition of modules within supernet's computation graph. This work explores the pathways proposal from the lens of statistical learning: We first develop novel generalization bounds for empirical risk minimization problems learning multiple tasks over multiple paths (Multipath MTL). In conjunction, we formalize the benefits of resulting multipath representation when adapting to new downstream tasks. Our bounds are expressed in terms of Gaussian complexity, lead to tangible guarantees for the class of linear representations, and provide novel insights into the quality and benefits of a multipath representation. When computation graph is a tree, Multipath MTL hierarchically clusters the tasks and builds cluster-specific representations. We provide further discussion and experiments for hierarchical MTL and rigorously identify the conditions under which Multipath MTL is provably superior to traditional MTL approaches with shallow supernets.
翻译:在多任务数量庞大的情况下,构建可迁移的表征是实现样本高效智能系统的关键需求。多任务学习(MTL)中的传统思路是构建跨任务的共享表征,然后通过调整最后几层来适配新任务。相比使用单一通用表征,构建任务特有表征是一种更理想的改进方案。为此,近期提出的PathNet/muNet架构将每个任务表示为更大超网络内部的路径。由路径诱导的子网络可视为超网络计算图中模块组合而成的任务特有表征。本文从统计学习视角探究路径方案:首先,我们针对多路径多任务学习(Multipath MTL)中的经验风险最小化问题,提出了新的泛化界。与此同时,我们形式化地阐述了当适配下游新任务时,所得多路径表征的优势。我们的泛化界用高斯复杂度表示,为线性表征类提供了具体的保证,并揭示了多路径表征质量及优势的新见解。当计算图为树结构时,Multipath MTL对任务进行层次聚类并构建聚类特有表征。我们进一步讨论了层次MTL并进行了实验,严格证明了在浅层超网络条件下Multipath MTL显著优于传统MTL方法的条件。