Conversational AI is increasingly used for advice, interpretation, reassurance, and decision support in contexts where users may be vulnerable, uncertain, or dependent on the system's apparent competence. Existing alignment work often focuses on model objectives, preference optimization, or output correctness. Yet, many harms arise through interaction: how systems frame authority, express uncertainty, simulate empathy, support reasoning, and make boundaries legible. This paper introduces the Layered Cognitive Alignment Model (LCAM), a conceptual and normative framework for diagnosing interac-tional alignment failures in conversational AI. LCAM defines alignment as a calibrated fit among system behavior, user goals, task demands, and normative context. It distinguishes five layers of fit: perceptual, semantic, affective, cognitive, and ethical, and two diagnostic polarities of misalignment: underfit and overreach. We apply LCAM to a published LLM counseling example, showing how an apparently supportive response can reinforce harmful beliefs, simulate inappropriate care, and obscure role boundaries. By translating conversational failures into audit and governance questions concerning over-reliance, false intimacy, autonomy erosion, boundary confusion, and inappropriate trust, LCAM offers a theoretical and normative lens for evaluating conversational AI beyond accuracy, helpfulness, or trust.


翻译:对话式AI越来越多地用于用户在脆弱、不确定或依赖系统表面能力的情境中提供建议、解读、安慰和决策支持。现有的对齐工作常聚焦于模型目标、偏好优化或输出正确性。然而,许多危害源于交互本身:系统如何构建权威、表达不确定性、模拟共情、支持推理以及界定边界。本文引入了分层认知对齐模型(LCAM),这是一个概念性和规范性框架,用于诊断对话式AI中的交互对齐失败。LCAM将对齐定义为系统行为、用户目标、任务需求和规范语境之间的校准匹配。它区分了五个匹配层面:感知层、语义层、情感层、认知层和伦理层,以及两种失调的诊断极性:欠匹配和过度干预。我们将LCAM应用于一个已发表的LLM心理咨询实例,展示了一个看似支持性的回应如何强化有害信念、模拟不恰当的关怀并模糊角色边界。通过将对话失败转化为关于过度依赖、虚假亲密、自主性侵蚀、边界混淆和不恰当信任的审计与治理问题,LCAM为在准确性、有用性或信任之外评估对话式AI提供了理论性和规范性的视角。

0
下载
关闭预览

相关内容

112页《人工智能对齐:全面性综述》中文版
专知会员服务
160+阅读 · 2024年2月1日
NeurIPS2020最新《深度对话人工智能》教程,130页ppt
专知会员服务
43+阅读 · 2020年12月10日
对话系统近期进展
专知
37+阅读 · 2019年3月23日
知识在检索式对话系统的应用
微信AI
32+阅读 · 2018年9月20日
干货篇|百度UNIT对话系统核心技术解析
InfoQ
23+阅读 · 2018年9月20日
最新人机对话系统简略综述
专知
26+阅读 · 2018年3月10日
一文读懂智能对话系统
数据派THU
16+阅读 · 2018年1月27日
基于 rasa 搭建中文对话系统 | 公开课
AI研习社
16+阅读 · 2018年1月12日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
VIP会员
最新内容
驱动军事决策变革的顶尖人工智能指挥系统
专知会员服务
7+阅读 · 8月11日
非对称防御中的自组织临界性:俄乌战争
专知会员服务
10+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
13+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
12+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
13+阅读 · 8月8日
相关VIP内容
112页《人工智能对齐:全面性综述》中文版
专知会员服务
160+阅读 · 2024年2月1日
NeurIPS2020最新《深度对话人工智能》教程,130页ppt
专知会员服务
43+阅读 · 2020年12月10日
相关资讯
对话系统近期进展
专知
37+阅读 · 2019年3月23日
知识在检索式对话系统的应用
微信AI
32+阅读 · 2018年9月20日
干货篇|百度UNIT对话系统核心技术解析
InfoQ
23+阅读 · 2018年9月20日
最新人机对话系统简略综述
专知
26+阅读 · 2018年3月10日
一文读懂智能对话系统
数据派THU
16+阅读 · 2018年1月27日
基于 rasa 搭建中文对话系统 | 公开课
AI研习社
16+阅读 · 2018年1月12日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员