Although disentangled representations are often said to be beneficial for downstream tasks, current empirical and theoretical understanding is limited. In this work, we provide evidence that disentangled representations coupled with sparse base-predictors improve generalization. In the context of multi-task learning, we prove a new identifiability result that provides conditions under which maximally sparse base-predictors yield disentangled representations. Motivated by this theoretical result, we propose a practical approach to learn disentangled representations based on a sparsity-promoting bi-level optimization problem. Finally, we explore a meta-learning version of this algorithm based on group Lasso multiclass SVM base-predictors, for which we derive a tractable dual formulation. It obtains competitive results on standard few-shot classification benchmarks, while each task is using only a fraction of the learned representations.
翻译:尽管解缠表征常被认为对下游任务有益,但目前对其在实证和理论层面的理解仍十分有限。本研究提供了证据表明,解缠表征结合稀疏基预测器可提升泛化性能。在多任务学习背景下,我们证明了一个新的可辨识性结果,该结果给出了最大稀疏基预测器产生解缠表征的条件。受此理论结果的启发,我们提出了一种基于稀疏促进双层优化问题的实用方法,用于学习解缠表征。最后,我们探索了该算法的一个元学习版本,其基于组Lasso多类SVM基预测器,并推导出其易于处理的对偶形式。该方法在标准小样本分类基准上取得了具有竞争力的结果,而每个任务仅使用了学习表征的一小部分。