This paper examines various methods of computing uncertainty and diversity for active learning in genetic programming. We found that the model population in genetic programming can be exploited to select informative training data points by using a model ensemble combined with an uncertainty metric. We explored several uncertainty metrics and found that differential entropy performed the best. We also compared two data diversity metrics and found that correlation as a diversity metric performs better than minimum Euclidean distance, although there are some drawbacks that prevent correlation from being used on all problems. Finally, we combined uncertainty and diversity using a Pareto optimization approach to allow both to be considered in a balanced way to guide the selection of informative and unique data points for training.
翻译:本文研究了遗传编程中主动学习的各种不确定性及多样性计算方法。我们发现,通过结合模型集成与不确定性度量,可以利用遗传编程中的模型群体来选择具有信息价值的训练数据点。我们探索了多种不确定性度量指标,发现差分熵表现最佳。同时,我们比较了两种数据多样性度量方法,结果表明相关性作为多样性度量指标优于最小欧氏距离,但存在某些局限性使得相关性无法应用于所有问题。最后,我们采用帕累托优化方法将不确定性与多样性相结合,使两者得以均衡考量,从而引导选择具有信息价值且独特的数据点用于训练。