Hyperparameters play a critical role in machine learning. Hyperparameter tuning can make the difference between state-of-the-art and poor prediction performance for any algorithm, but it is particularly challenging for structure learning due to its unsupervised nature. As a result, hyperparameter tuning is often neglected in favour of using the default values provided by a particular implementation of an algorithm. While there have been numerous studies on performance evaluation of causal discovery algorithms, how hyperparameters affect individual algorithms, as well as the choice of the best algorithm for a specific problem, has not been studied in depth before. This work addresses this gap by investigating the influence of hyperparameters on causal structure learning tasks. Specifically, we perform an empirical evaluation of hyperparameter selection for some seminal learning algorithms on datasets of varying levels of complexity. We find that, while the choice of algorithm remains crucial to obtaining state-of-the-art performance, hyperparameter selection in ensemble settings strongly influences the choice of algorithm, in that a poor choice of hyperparameters can lead to analysts using algorithms which do not give state-of-the-art performance for their data.
翻译:超参数在机器学习中扮演着关键角色。超参数调优可能决定任何算法的性能是达到顶尖水平还是表现不佳,但对于结构学习而言,由于其无监督特性,这一任务尤为棘手。因此,超参数调优常被忽视,转而使用特定算法实现提供的默认参数。尽管已有大量关于因果发现算法性能评估的研究,但超参数如何影响单个算法以及如何针对特定问题选择最佳算法,此前尚未得到深入探讨。本文通过研究超参数对因果结构学习任务的影响来填补这一空白。具体而言,我们对几种经典学习算法在不同复杂度数据集上的超参数选择进行了实证评估。研究发现,尽管算法选择对获得顶尖性能仍至关重要,但集成设置中的超参数选择会强烈影响算法选择:超参数选择不当可能导致分析人员使用对其数据无法提供顶尖性能的算法。