A soft tree is an actively studied variant of a decision tree that updates splitting rules using the gradient method. Although soft trees can take various architectures, their impact is not theoretically well known. In this paper, we formulate and analyze the Neural Tangent Kernel (NTK) induced by soft tree ensembles for arbitrary tree architectures. This kernel leads to the remarkable finding that only the number of leaves at each depth is relevant for the tree architecture in ensemble learning with an infinite number of trees. In other words, if the number of leaves at each depth is fixed, the training behavior in function space and the generalization performance are exactly the same across different tree architectures, even if they are not isomorphic. We also show that the NTK of asymmetric trees like decision lists does not degenerate when they get infinitely deep. This is in contrast to the perfect binary trees, whose NTK is known to degenerate and leads to worse generalization performance for deeper trees.
翻译:软树是决策树的一个活跃变体,通过梯度方法更新分裂规则。尽管软树可采用多种架构,但其影响在理论上尚不明确。本文针对任意树架构,建立并分析了软树集成所诱导的神经正切核(NTK)。该核揭示了重要发现:在无穷树数量的集成学习中,仅每层叶节点数量对树架构具有相关性。换言之,若各层叶节点数量固定,不同树架构(即使非同构)在函数空间中的训练行为与泛化性能完全一致。本文还证明,当决策列表等非对称树趋于无穷深时,其NTK不会退化。这与完美二叉树形成对比——后者的NTK已知会退化,并导致更深树产生更差的泛化性能。