Learning on big data brings success for artificial intelligence (AI), but the annotation and training costs are expensive. In future, learning on small data that approximates the generalization ability of big data is one of the ultimate purposes of AI, which requires machines to recognize objectives and scenarios relying on small data as humans. A series of learning topics is going on this way such as active learning and few-shot learning. However, there are few theoretical guarantees for their generalization performance. Moreover, most of their settings are passive, that is, the label distribution is explicitly controlled by finite training resources from known distributions. This survey follows the agnostic active sampling theory under a PAC (Probably Approximately Correct) framework to analyze the generalization error and label complexity of learning on small data in model-agnostic supervised and unsupervised fashion. Considering multiple learning communities could produce small data representation and related topics have been well surveyed, we thus subjoin novel geometric representation perspectives for small data: the Euclidean and non-Euclidean (hyperbolic) mean, where the optimization solutions including the Euclidean gradients, non-Euclidean gradients, and Stein gradient are presented and discussed. Later, multiple learning communities that may be improved by learning on small data are summarized, which yield data-efficient representations, such as transfer learning, contrastive learning, graph representation learning. Meanwhile, we find that the meta-learning may provide effective parameter update policies for learning on small data. Then, we explore multiple challenging scenarios for small data, such as the weak supervision and multi-label. Finally, multiple data applications that may benefit from efficient small data representation are surveyed.
翻译:大数据学习为人工智能带来了成功,但标注和训练成本高昂。未来,能够近似大数据泛化能力的小数据学习是人工智能的终极目标之一,这要求机器像人类一样仅依赖小数据识别目标和场景。沿着这一方向,主动学习、小样本学习等一系列学习课题正在推进。然而,这些方法的泛化性能鲜有理论保障。此外,其设置大多是被动的,即标签分布由已知分布中的有限训练资源显式控制。本综述遵循PAC(可能近似正确)框架下的不可知主动采样理论,分析小数据学习在模型无关的监督与无监督方式下的泛化误差和标签复杂度。鉴于多个学习社区已产出小数据表征且相关主题已得到充分综述,本文补充了新颖的几何表征视角:欧几里得均值与非欧几里得(双曲)均值,其中讨论了包括欧几里得梯度、非欧几里得梯度与斯坦因梯度在内的优化方案。随后,总结了可能通过小数据学习改进的多个学习社区(如迁移学习、对比学习、图表示学习),它们能产生数据高效的表示。同时,我们发现元学习可能为小数据学习提供有效的参数更新策略。接着,探讨了小数据面临的多个挑战性场景,例如弱监督与多标签。最后,综述了可能受益于高效小数据表示的多类数据应用。