Structured tabular data is a fundamental data type in numerous fields, and the capacity to reason over tables is crucial for answering questions and validating hypotheses. However, constructing labeled data for complex reasoning tasks is labor intensive, and the quantity of annotated data remains insufficient to support the intricate demands of real-world applications. To address the insufficient annotation challenge, we present a self-training framework for unsupervised complex tabular reasoning (UCTR-ST) by generating diverse synthetic data with complex logic. Specifically, UCTR-ST incorporates several essential techniques: we aggregate diverse programs and execute them on tables based on a "Program-Management" component, and we bridge the gap between programs and text with a powerful "Program-Transformation" module that generates natural language sentences with complex logic. Furthermore, we optimize the procedure using a "Table-Text Manipulator" to handle joint table-text reasoning scenarios. The entire framework utilizes self-training techniques to leverage the unlabeled training data, which results in significant performance improvements when tested on real-world data. Experimental results demonstrate that UCTRST achieves above 90% of the supervised model performance on different tasks and domains, reducing the dependence on manual annotation. Additionally, our approach can serve as a data augmentation technique, significantly boosting the performance of supervised models in low-resourced domains.
翻译:结构化表格数据是众多领域中的基础数据类型,对表格进行推理的能力对于回答问题和验证假设至关重要。然而,为复杂推理任务构建标注数据需要大量人力,且标注数据的数量仍不足以满足现实应用中复杂场景的需求。为解决标注不足的挑战,我们提出了一种用于无监督复杂表格推理的自训练框架(UCTR-ST),通过生成具有复杂逻辑的多样化合成数据来实现。具体而言,UCTR-ST整合了多项关键技术:我们基于“程序管理”组件聚合多样化程序并在表格上执行它们;同时,我们通过一个强大的“程序转换”模块弥合程序与文本之间的鸿沟,该模块能够生成具有复杂逻辑的自然语言句子。此外,我们利用“表格-文本操纵器”优化处理流程,以应对联合表格-文本推理场景。整个框架采用自训练技术以充分利用未标注的训练数据,在真实数据测试中实现了显著的性能提升。实验结果表明,UCTR-ST在不同任务和领域上均能达到有监督模型90%以上的性能,降低了对人工标注的依赖。此外,我们的方法可作为数据增强技术,显著提升低资源领域中有监督模型的性能。