Standard evaluation practices assume that large language model (LLM) outputs are stable when prompts are embedded in contextually equivalent discourses. Here, we test this assumption in the setting of gender inference. Using a controlled pronoun selection task, we introduce minimal, theoretically uninformative discourse context and find that this induces large, systematic shifts in model outputs. Correlations with cultural gender stereotypes, present in decontextualized settings, weaken or disappear once context is introduced, while theoretically irrelevant features, such as the gender of a pronoun for an unrelated referent, become the most informative predictors of model behavior. A Contextuality-by-Default analysis reveals that, in 19--52\% of cases across models, this dependence persists after accounting for all marginal effects of context on individual outputs and cannot be attributed to simple pronoun repetition. These findings show that LLM outputs violate contextual invariance even under near-identical syntactic formulations, with implications for bias benchmarking and deployment in high-stakes settings.
翻译:标准评估实践假设,当提示嵌入到上下文等价的语篇中时,大语言模型(LLM)的输出是稳定的。在此,我们在性别推断的设定下检验这一假设。通过使用受控代词选择任务,我们引入了微小的、理论上无信息的语篇上下文,并发现这会导致模型输出发生大规模系统性偏移。与去语境化设定中存在的文化性别刻板印象的相关性,在引入上下文后会减弱或消失,而理论上无关的特征(例如,无关指代对象的代词性别)则成为模型行为最具信息量的预测因子。一项"默认上下文性"分析表明,在跨模型的19%至52%的案例中,在考虑上下文对单个输出的所有边际效应后,这种依赖性仍然存在,且不能归因于简单的代词重复。这些发现表明,即使在句法结构近乎相同的情况下,LLM输出也违反了上下文不变性,这对偏见基准测试及高风险场景中的部署具有重要意义。