We study attribute control in language models through the method of Causal Average Treatment Effect (Causal ATE). Existing methods for the attribute control task in Language Models (LMs) check for the co-occurrence of words in a sentence with the attribute of interest, and control for them. However, spurious correlation of the words with the attribute in the training dataset, can cause models to hallucinate the presence of the attribute when presented with the spurious correlate during inference. We show that the simple perturbation-based method of Causal ATE removes this unintended effect. Additionally, we offer a theoretical foundation for investigating Causal ATE in the classification task, and prove that it reduces the number of false positives -- thereby mitigating the issue of unintended bias. Specifically, we ground it in the problem of toxicity mitigation, where a significant challenge lies in the inadvertent bias that often emerges towards protected groups post detoxification. We show that this unintended bias can be solved by the use of the Causal ATE metric.
翻译:我们通过因果平均处理效应(Causal ATE)方法研究语言模型中的属性控制。现有语言模型属性控制任务的方法会检查句子中词语与目标属性的共现情况,并据此进行控制。然而,训练数据集中词语与属性之间的虚假相关性会导致模型在推理时遇到虚假相关词时产生属性幻觉。我们证明,基于简单扰动方法的因果平均处理效应能够消除这种意外效应。此外,我们为分类任务中因果平均处理效应的研究提供了理论基础,并证明其能够降低假阳性数量——从而缓解意外偏差问题。具体而言,我们将其应用于毒性缓解问题中,其核心挑战在于去毒化后往往出现针对受保护群体的无意识偏差。我们证明,这种无意识偏差可通过使用因果平均处理效应指标得以解决。