Standard evaluation practices assume that large language model (LLM) outputs are stable under contextually equivalent formulations of a task. Here, we test this assumption in the setting of gender inference. Using a controlled pronoun selection task, we introduce minimal, theoretically uninformative discourse context and find that this induces large, systematic shifts in model outputs. Correlations with cultural gender stereotypes, present in decontextualized settings, weaken or disappear once context is introduced, while theoretically irrelevant features, such as the gender of a pronoun for an unrelated referent, become the most informative predictors of model behaviour. A Contextuality-by-Default analysis reveals that, in 19--52\% of cases across models, this dependence persists after accounting for all marginal effects of context on individual outputs and cannot be attributed to simple pronoun repetition. These findings show that LLM outputs violate contextual invariance even under near-identical syntactic formulations, with implications for bias benchmarking and deployment in high-stakes settings.
翻译:标准评估实践假定,大语言模型(LLM)在任务具有上下文等价表述时,其输出保持稳定。本文在性别推断情境下检验了这一假设。通过一项受控代词选择任务,我们引入最小化且理论上无信息量的语篇上下文,发现这会导致模型输出出现大规模系统性偏移。去语境化情境中存在的与文化性别刻板印象的关联,在引入上下文后减弱或消失;而理论无关特征(如无关所指代词性别)则成为模型行为最具信息量的预测因子。采用“默认语境性”分析表明,在模型间19%-52%的案例中,这种依赖性在控制上下文对个体输出的所有边际效应后依然存在,且不能归因于简单的代词重复。这些发现表明,即便在句法表述近乎一致的情况下,LLM的输出仍违反上下文不变性,这对偏差评测及高风险场景下的部署具有重要启示。