Emotional intelligence significantly impacts our daily behaviors and interactions. Although Large Language Models (LLMs) are increasingly viewed as a stride toward artificial general intelligence, exhibiting impressive performance in numerous tasks, it is still uncertain if LLMs can genuinely grasp psychological emotional stimuli. Understanding and responding to emotional cues gives humans a distinct advantage in problem-solving. In this paper, we take the first step towards exploring the ability of LLMs to understand emotional stimuli. To this end, we first conduct automatic experiments on 45 tasks using various LLMs, including Flan-T5-Large, Vicuna, Llama 2, BLOOM, ChatGPT, and GPT-4. Our tasks span deterministic and generative applications that represent comprehensive evaluation scenarios. Our automatic experiments show that LLMs have a grasp of emotional intelligence, and their performance can be improved with emotional prompts (which we call "EmotionPrompt" that combines the original prompt with emotional stimuli), e.g., 8.00% relative performance improvement in Instruction Induction and 115% in BIG-Bench. In addition to those deterministic tasks that can be automatically evaluated using existing metrics, we conducted a human study with 106 participants to assess the quality of generative tasks using both vanilla and emotional prompts. Our human study results demonstrate that EmotionPrompt significantly boosts the performance of generative tasks (10.9% average improvement in terms of performance, truthfulness, and responsibility metrics). We provide an in-depth discussion regarding why EmotionPrompt works for LLMs and the factors that may influence its performance. We posit that EmotionPrompt heralds a novel avenue for exploring interdisciplinary knowledge for human-LLMs interaction.
翻译:情感智能对我们的日常行为和互动有着显著影响。尽管大型语言模型(LLMs)日益被视为迈向人工通用智能的一步,并在众多任务中展现出令人瞩目的性能,但LLMs是否真正理解心理情感刺激仍不确定。理解和回应情感线索使人类在解决问题中具备独特优势。本文首次探索了LLMs理解情感刺激的能力。为此,我们首先使用多种LLMs(包括Flan-T5-Large、Vicuna、Llama 2、BLOOM、ChatGPT和GPT-4)在45项任务上进行了自动化实验。这些任务涵盖确定性任务和生成式任务,代表了全面的评估场景。我们的自动化实验表明,LLMs具备情感智能的把握能力,且其性能可通过情感提示(我们称之为"EmotionPrompt"——将原始提示与情感刺激相结合)得到提升,例如在Instruction Induction任务中相对性能提升8.00%,在BIG-Bench中提升115%。除这些可使用现有指标自动评估的确定性任务外,我们招募了106名参与者进行人工研究,评估使用原始提示和情感提示的生成式任务质量。人工研究结果表明,EmotionPrompt显著提升了生成式任务的性能(在性能、真实性、责任性指标上平均提升10.9%)。我们深入讨论了EmotionPrompt为何对LLMs有效以及可能影响其性能的因素。我们认为,EmotionPrompt为探索人机交互的跨学科知识开辟了一条新途径。