As Large Language Models (LLMs) become embedded in everyday communication, capturing regional linguistic variation is essential for reliable and equitable language use. In Portuguese, European (pt-PT) and Brazilian (pt-BR) varieties remain unevenly represented, with pt-BR dominating in data quantity, while LLM preference for Portuguese variants remains underexplored. To address this gap, we introduce P3B3, an expert-curated language variety agnostic benchmark of conversational prompts, along with an evaluation framework for measuring variety bias and controllability. Experiments on several models show that most LLMs exhibit a strong bias toward pt-BR, with variation in controllability across models. These results highlight the need for more balanced multilingual representation across language varieties.
翻译:随着大型语言模型(LLM)深度融入日常交流,捕捉区域语言变体对于实现可靠且公平的语言使用至关重要。在葡萄牙语中,欧洲(pt-PT)与巴西(pt-BR)变体的代表性仍不均衡:数据量上巴式葡萄牙语占据主导地位,而LLM对葡萄牙语变体的偏好尚未得到充分探究。为填补这一空白,我们提出P3B3——一个由专家策展、无关语言变体的会话提示基准,并配套构建了评估变体偏差及可控性的框架。在多个模型上的实验表明,绝大多数LLM对巴式葡萄牙语存在显著偏好,且各模型间的可控性存在差异。这些结果凸显了针对不同语言变体实现更均衡多语言表征的必要性。