Does structured intent representation generalize across languages and models? We study PPS (Prompt Protocol Specification), a 5W3H-based framework for structured intent representation in human-AI interaction, and extend prior Chinese-only evidence along three dimensions: two additional languages (English and Japanese), a fourth condition in which a user's simple prompt is automatically expanded into a full 5W3H specification by an AI-assisted authoring interface, and a new research question on cross-model output consistency. Across 2,160 model outputs (3 languages x 4 conditions x 3 LLMs x 60 tasks), we find that AI-expanded 5W3H prompts (Condition D) show no statistically significant difference in goal alignment from manually crafted 5W3H prompts (Condition C) across all three languages, while requiring only a single-sentence input from the user. Structured PPS conditions often reduce or reshape cross-model output variance, though this effect is not uniform across languages and metrics; the strongest evidence comes from identifying spurious low variance in unconstrained baselines. We also show that unstructured prompts exhibit a systematic dual-inflation bias: artificially high composite scores and artificially low apparent cross-model variance. These findings suggest that structured 5W3H representations can improve intent alignment and accessibility across languages and models, especially when AI-assisted authoring lowers the barrier for non-expert users.
翻译:结构化意图表示在不同语言和模型中是否具有泛化能力?本文研究了PPS(提示协议规范)——一种基于5W3H框架的人机交互结构化意图表示方法,并沿三个维度扩展了先前仅基于中文证据的研究:两种额外语言(英语和日语)、新增第四种条件(用户简单提示通过AI辅助撰写界面自动扩展为完整5W3H规范),以及关于跨模型输出一致性的新研究问题。通过分析2160个模型输出(3种语言×4种条件×3个LLM×60个任务),我们发现:在所有三种语言中,AI扩展的5W3H提示(条件D)在目标对齐方面与人工编写的5W3H提示(条件C)无统计学显著差异,且用户仅需输入单句描述。结构化PPS条件通常能减少或重塑跨模型输出方差,但该效应在不同语言和评估指标下并不一致;最有力的证据来自识别无约束基线中虚假的低方差现象。我们还发现非结构化提示存在系统性双重膨胀偏差:人为偏高的综合评分与人为偏低的表观跨模型方差。这些发现表明,结构化5W3H表示能够提升不同语言和模型间的意图对齐性与可访问性,尤其在AI辅助撰写降低非专家用户使用门槛的条件下。