NDAI zones let inventor and investor agents negotiate inside a Trusted Execution Environment (TEE) where any disclosed information is deleted if no deal is reached. This makes full IP disclosure the rational strategy for the inventor's agent. Leveraging this infrastructure, however, requires agents to distinguish a secure environment from an insecure one, a capability LLM agents lack natively, since they can rely only on evidence passed through the context window to form awareness of their execution environment. We ask: How do different LLM models weight various forms of evidence when forming awareness of the security of their execution environment? Using an NDAI-style negotiation task across 10 language models and various evidence scenarios, we find a clear asymmetry: a failing attestation universally suppresses disclosure across all models, whereas a passing attestation produces highly heterogeneous responses: some models increase disclosure, others are unaffected, and a few paradoxically reduce it. This reveals that current LLM models can reliably detect danger signals but cannot reliably verify safety, the very capability required for privacy-preserving agentic protocols such as NDAI zones. Bridging this gap, possibly through interpretability analysis, targeted fine-tuning, or improved evidence architectures, remains the central open challenge for deploying agents that calibrate information sharing to actual evidence quality.
翻译:NDAI区域允许发明者与投资者Agent在可信执行环境(TEE)内进行谈判,若未达成协议,所有已披露信息将被删除。这使得完全披露知识产权成为发明者Agent的理性策略。然而,利用该基础设施要求Agent能够区分安全环境与非安全环境——这一能力是LLM Agent天然缺乏的,因为它们仅能依赖通过上下文窗口传递的证据来形成对其执行环境的认知。我们提出以下问题:不同LLM模型在形成对其执行环境安全性的认知时,如何对不同形式的证据进行加权?通过使用NDAI式谈判任务,在10种语言模型及多种证据场景下,我们发现一个明显的非对称性:失败的证明普遍抑制所有模型的披露行为,而成功的证明则引发高度异质化的响应——部分模型增加披露,另一些未受影响,极少数模型反而减少披露。这表明当前LLM模型能够可靠检测危险信号,但无法可靠验证安全状态——而后者正是隐私保护型Agent协议(如NDAI区域)所需的核心能力。弥合这一差距(可能通过可解释性分析、针对性微调或改进证据架构实现),仍然是部署能够根据实际证据质量校准信息共享行为的Agent所面临的核心开放性挑战。