This paper explores the response of Large Language Models (LLMs) to user prompts with different degrees of politeness and impoliteness. The Politeness Theory by Brown and Levinson and the Impoliteness Framework by Culpeper form the basis of experiments conducted across three languages (English, Hindi, Spanish), five models (Gemini-Pro, GPT-4o Mini, Claude 3.7 Sonnet, DeepSeek-Chat, and Llama 3), and three interaction histories between users (raw, polite, and impolite). Our sample consists of 22,500 pairs of prompts and responses of various types, evaluated across five levels of politeness using an eight-factor assessment framework: coherence, clarity, depth, responsiveness, context retention, toxicity, conciseness, and readability. The findings show that model performance is highly influenced by tone, dialogue history, and language. While polite prompts enhance the average response quality by up to ~11% and impolite tones worsen it, these effects are neither consistent nor universal across languages and models. English is best served by courteous or direct tones, Hindi by deferential and indirect tones, and Spanish by assertive tones. Among the models, Llama is the most tone-sensitive (11.5% range), whereas GPT is more robust to adversarial tone. These results indicate that politeness is a quantifiable computational variable that affects LLM behaviour, though its impact is language- and model-dependent rather than universal. To support reproducibility and future work, we additionally release PLUM (Politeness Levels in Utterances, Multilingual), a publicly available corpus of 1,500 human-validated prompts across three languages and five politeness categories, and provide a formal supplementary analysis of six falsifiable hypotheses derived from politeness theory, empirically assessed against the dataset.
翻译:本文探究大语言模型对用户不同礼貌与不礼貌程度提示词的响应。基于Brown和Levinson的礼貌理论与Culpeper的不礼貌理论框架,我们开展涉及三种语言(英语、印地语、西班牙语)、五个模型(Gemini-Pro、GPT-4o Mini、Claude 3.7 Sonnet、DeepSeek-Chat以及Llama 3)和三种用户交互历史(原始、礼貌、不礼貌)的实验。样本包含22,500组各类提示词与响应配对,采用八因子评估框架(连贯性、清晰度、深度、响应性、上下文保持、毒性、简洁性和可读性)按五个礼貌层级进行评价。研究发现模型性能深受语气、对话历史及语言影响。尽管礼貌提示词可将平均响应质量提升约11%,不礼貌语气会降低质量,但这些效应在语言和模型间既不一致也不具有普遍性。英语最适合礼貌或直接语气,印地语偏好恭敬和间接语气,西班牙语则倾向坚定语气。模型方面,Llama对语气最敏感(波动范围达11.5%),而GPT对对抗性语气更具鲁棒性。结果表明礼貌是可量化影响大语言模型行为的计算变量,但其影响具有语言和模型依赖性而非普遍性。为促进可重复性与后续研究,我们同时发布PLUM(多语言话语礼貌层级)语料库——包含三种语言、五个礼貌类别共1,500条经人工验证的提示词公开数据集,并基于该数据集对源自礼貌理论的六项可证伪假设进行规范的补充分析。