In some machine learning applications the availability of labeled instances for supervised classification is limited while unlabeled instances are abundant. Semi-supervised learning algorithms deal with these scenarios and attempt to exploit the information contained in the unlabeled examples. In this paper, we address the question of how to evolve neural networks for semi-supervised problems. We introduce neuroevolutionary approaches that exploit unlabeled instances by using neuron coverage metrics computed on the neural network architecture encoded by each candidate solution. Neuron coverage metrics resemble code coverage metrics used to test software, but are oriented to quantify how the different neural network components are covered by test instances. In our neuroevolutionary approach, we define fitness functions that combine classification accuracy computed on labeled examples and neuron coverage metrics evaluated using unlabeled examples. We assess the impact of these functions on semi-supervised problems with a varying amount of labeled instances. Our results show that the use of neuron coverage metrics helps neuroevolution to become less sensitive to the scarcity of labeled data, and can lead in some cases to a more robust generalization of the learned classifiers.
翻译:在某些机器学习应用中,用于监督分类的标注实例有限,而未标注实例却十分丰富。半监督学习算法处理此类场景,并试图利用未标注样本中包含的信息。本文探讨如何针对半监督问题进化神经网络。我们引入了利用未标注实例的神经进化方法,该方法通过计算每个候选解编码的神经网络架构上的神经元覆盖度指标来实现。神经元覆盖度指标类似于用于测试软件的代码覆盖度指标,但侧重于量化测试实例对神经网络不同组件的覆盖程度。在我们的神经进化方法中,我们定义了结合标注实例分类精度和未标注实例神经元覆盖度指标的适应度函数。我们评估了这些函数在不同标注实例数量的半监督问题上的影响。结果表明,使用神经元覆盖度指标有助于神经进化对标注数据稀缺的敏感性降低,并且在某些情况下可导致学习分类器更鲁棒的泛化能力。