The transferability of deep neural networks (DNNs) has made significant progress in image and language processing. However, due to the heterogeneity among tables, such DNN bonus is still far from being well exploited on tabular data prediction (e.g., regression or classification tasks). Condensing knowledge from diverse domains, language models (LMs) possess the capability to comprehend feature names from various tables, potentially serving as versatile learners in transferring knowledge across distinct tables and diverse prediction tasks, but their discrete text representation space is inherently incompatible with numerical feature values in tables. In this paper, we present TP-BERTa, a specifically pre-trained LM for tabular data prediction. Concretely, a novel relative magnitude tokenization converts scalar numerical feature values to finely discrete, high-dimensional tokens, and an intra-feature attention approach integrates feature values with the corresponding feature names. Comprehensive experiments demonstrate that our pre-trained TP-BERTa leads the performance among tabular DNNs and is competitive with Gradient Boosted Decision Tree models in typical tabular data regime.
翻译:深度神经网络(DNN)的可迁移性在图像和语言处理领域取得了显著进展。然而,由于表格数据的异质性,这类DNN优势在表格数据预测(如回归或分类任务)中尚未得到充分利用。语言模型(LM)汇聚了来自不同领域的知识,能够理解各类表格中的特征名称,有望成为跨不同表格和多样化预测任务进行知识迁移的多功能学习器,但其离散文本表示空间与表格中的数值特征值天然不兼容。本文提出TP-BERTa——一种专为表格数据预测而预训练的语言模型。具体而言,一种新颖的相对幅度分词方法将标量数值特征值转化为精细离散的高维词元,而特征内注意力机制则将特征值与对应的特征名称相融合。综合实验表明,我们预训练的TP-BERTa在表格DNN中表现领先,并在典型表格数据任务中与梯度提升决策树模型不相上下。