Numerous works are proposed to align large language models (LLMs) with human intents to better fulfill instructions, ensuring they are trustful and helpful. Nevertheless, some human instructions are often malicious or misleading and following them will lead to untruthful and unsafe responses. Previous work rarely focused on understanding how LLMs manage instructions based on counterfactual premises, referred to here as \textit{inductive instructions}, which may stem from users' false beliefs or malicious intents. In this paper, we aim to reveal the behaviors of LLMs towards \textit{inductive instructions} and enhance their truthfulness and helpfulness accordingly. Specifically, we first introduce a benchmark of \underline{\textbf{Indu}}ctive {In\underline{\textbf{st}}ruct}ions (\textsc{\textbf{INDust}}), where the false knowledge is incorporated into instructions in multiple different styles. After extensive human and automatic evaluations, we uncovered a universal vulnerability among LLMs in processing inductive instructions. Additionally, we identified that different inductive styles affect the models' ability to identify the same underlying errors, and the complexity of the underlying assumptions also influences the model's performance. Motivated by these results, we propose \textsc{Dual-critique} prompting to improve LLM robustness against inductive instructions. Our experiments demonstrate that \textsc{Dual-critique} prompting significantly bolsters the robustness of a diverse array of LLMs, even when confronted with varying degrees of inductive instruction complexity and differing inductive styles.
翻译:现有大量研究致力于对齐大语言模型(LLMs)与人类意图,使其能更好地遵循指令,从而确保其可信性与有用性。然而,部分人类指令常带有恶意或误导性,遵循此类指令将导致不真实、不安全的响应。以往工作鲜少关注模型如何处理基于反事实前提的指令(本文称为**归纳指令**),这类指令可能源自用户的错误认知或恶意意图。本文旨在揭示LLMs对归纳指令的行为特征,并据此提升其真实性与有用性。具体而言,我们首先构建了包含多种风格错误知识的归纳指令基准数据集(**INDust**)。通过大规模人工与自动评估,发现LLMs在处理归纳指令时存在普遍脆弱性。此外,我们还发现不同归纳风格会影响模型识别相同底层错误的能力,且底层假设的复杂度也会影响模型表现。基于上述发现,我们提出双批判提示方法(Dual-critique prompting)以增强LLMs对归纳指令的鲁棒性。实验表明,即使面对不同复杂度和风格的归纳指令,双批判提示方法仍能显著提升多种LLMs的鲁棒性。