Instruction tuning on a mixture of tasks has improved zero-shot capabilities in natural language processing (NLP). Nevertheless, existing methods often learn features that exhibit correlations between instruction-formatted samples and target labels, rather than causal relationships. Termed as ``spurious correlation'' in statistics, such a correlation may change drastically in a new task, making the effect from the learned features to be misleading. To this end, we develop a meta Structural Causal Model (meta-SCM) to integrate different NLP tasks under a single causal structure of the data. Specifically, the meta-SCM introduces multiple latent factors that represent properties of source context, only some of which causally influence the target labels for a specific task. The key idea is to learn task-required causal factors and only use those to make predictions for a given task. Theoretically, we prove the causal factor can be identified without mixing information from others. Guided by the identifiability, we propose a Structural Instruction Tuning (SIT) method to learn the task-required causal representations that can mimic the causal factors for each task. The utility of our approach is verified by improvements of zero-shot ability on a range of unseen datasets and tasks.
翻译:指令微调通过在混合任务上的训练提升了自然语言处理(NLP)中的零样本能力。然而,现有方法通常学习到的特征仅体现指令格式化样本与目标标签之间的相关性,而非因果关系。在统计学中,这种被称为“伪相关”的关系在新任务中可能发生剧烈变化,导致所学特征产生误导性影响。为此,我们提出了一种元结构因果模型(meta-SCM),将不同NLP任务整合到统一的数据因果结构中。具体而言,meta-SCM引入了多个表示源上下文属性的潜在因子,其中仅部分因子对特定任务的目标标签具有因果影响。其核心思想是学习任务所需的因果因子,并仅利用这些因子对给定任务进行预测。理论上,我们证明因果因子可以在不与其他信息混淆的情况下被识别。基于该可识别性,我们提出了一种结构指令微调(SIT)方法,学习能够模拟每个任务因果因子的任务所需因果表示。该方法在多个未见数据集和任务上的零样本能力提升验证了其有效性。