Large language model (LLM)-based persona agents are rapidly being adopted as scalable proxies for human participants across diverse domains. Yet there is no systematic method for verifying whether a persona agent's responses remain free of contradictions and factual inaccuracies throughout an interaction. A principle from interrogation methodology offers a lens: no matter how elaborate a fabricated identity, systematic interrogation will expose its contradictions. We apply this principle to propose PICon, an evaluation framework that probes persona agents through logically chained multi-turn questioning. PICon evaluates consistency along three core dimensions: internal consistency (freedom from self-contradiction), external consistency (alignment with real-world facts), and retest consistency (stability under repetition). Evaluating seven groups of persona agents alongside 63 real human participants, we find that even systems previously reported as highly consistent fail to meet the human baseline across all three dimensions, revealing contradictions and evasive responses under chained questioning. This work provides both a conceptual foundation and a practical methodology for evaluating persona agents before trusting them as substitutes for human participants. We provide the source code and an interactive demo at: https://kaist-edlab.github.io/picon/
翻译:基于大语言模型的人格智能体正快速被用作跨领域人类参与者的可扩展代理。然而,目前尚无系统化方法验证人格智能体的响应在交互过程中是否保持无矛盾且无事实错误。审讯方法论中的一条原则提供了视角:无论虚构身份多么精心构建,系统性审讯终将暴露其矛盾。我们应用该原则提出PICon,一种通过逻辑链式多轮提问探测人格智能体的评估框架。PICon沿三个核心维度评估一致性:内部一致性(无自我矛盾)、外部一致性(符合现实事实)与重测一致性(重复下的稳定性)。在对七组人格智能体及63名真实人类参与者进行评估后,我们发现即便先前报告为高度一致的系统,在所有三个维度上均未达到人类基线水平,并在链式提问下暴露出矛盾与回避性响应。本工作为在信任人格智能体替代人类参与者之前评估其一致性提供了概念基础与实践方法论。我们已在以下链接提供源代码与交互式演示:https://kaist-edlab.github.io/picon/