As large language models (LLMs) increasingly act as collaborative partners, human--AI alignment is often evaluated through explicit task success, accuracy, or reward optimization. Yet many collaborative settings depend on tacit understanding: whether an agent can align with a human's evaluative stance or representational priors without clear objectives, communication, or feedback. To study this capacity, we develop a spectrum-placement task inspired by the social party game Wavelength, in which humans and agents independently place concepts along subjective spectra. We operationalize the Tacit Understanding Index (TUX) as a pairwise measure of similarity between human and agent judgments, and evaluate it with 241 human participants and 200 profile-conditioned LLM agents across four models. We find that nearest human--agent pairs in trait space achieve significantly higher TUX, suggesting that tacit alignment is structured by person-level characteristics rather than random similarity. Regression analyses show that TUX becomes more explainable as predictor sets become richer, with individual traits, decision-making styles, and confidence improving over aggregate trait-distance baselines. These findings suggest that tacit understanding between humans and LLMs is measurable, while revealing the limits of profile-based conditioning for capturing deeper representational alignment.
翻译:随着大型语言模型(LLMs)日益成为协作伙伴,人机对齐通常通过显式任务成功、准确性或奖励优化来评估。然而,许多协作情境依赖于默契理解:智能体能否在没有明确目标、沟通或反馈的情况下,与人类的评价立场或表征先验保持一致。为研究这一能力,我们受社交派对游戏Wavelength启发,开发了一项谱系定位任务,让人类和智能体独立将概念置于主观谱系上。我们将默契理解指数(TUX)定义为人类与智能体判断之间成对相似度的度量,并基于241名人类参与者和200个基于用户画像条件的LLM智能体对四个模型进行评估。研究发现,特质空间中最接近的人机配对达到显著更高的TUX,表明默契对齐由个体层面特征结构决定,而非随机相似性。回归分析显示,随着预测变量集更丰富,TUX的可解释性增强,个体特质、决策风格和信心等变量相比聚合特质距离基线更具预测力。这些发现表明,人类与LLMs之间的默契理解是可测量的,同时揭示了基于画像的条件设置方法在捕捉深层表征对齐方面的局限性。