在多轮对话中引发特定行为 (Eliciting Behaviors in Multi-Turn Conversations) - 专知论文

会员服务 ·

0

多轮对话 · 交互 · 目标模型 · 在线 · 分析 ·

2025 年 12 月 29 日

Eliciting Behaviors in Multi-Turn Conversations

翻译：在多轮对话中引发特定行为

Jing Huang,Shujian Zhang,Lun Wang,Andrew Hard,Rajiv Mathews,John Lambert

Identifying specific and often complex behaviors from large language models (LLMs) in conversational settings is crucial for their evaluation. Recent work proposes novel techniques to find natural language prompts that induce specific behaviors from a target model, yet they are mainly studied in single-turn settings. In this work, we study behavior elicitation in the context of multi-turn conversations. We first offer an analytical framework that categorizes existing methods into three families based on their interactions with the target model: those that use only prior knowledge, those that use offline interactions, and those that learn from online interactions. We then introduce a generalized multi-turn formulation of the online method, unifying single-turn and multi-turn elicitation. We evaluate all three families of methods on automatically generating multi-turn test cases. We investigate the efficiency of these approaches by analyzing the trade-off between the query budget, i.e., the number of interactions with the target model, and the success rate, i.e., the discovery rate of behavior-eliciting inputs. We find that online methods can achieve an average success rate of 45/19/77% with just a few thousand queries over three tasks where static methods from existing multi-turn conversation benchmarks find few or even no failure cases. Our work highlights a novel application of behavior elicitation methods in multi-turn conversation evaluation and the need for the community to move towards dynamic benchmarks.

翻译：从大型语言模型（LLMs）在对话场景中识别特定且通常复杂的行为，对于其评估至关重要。近期研究提出了新颖技术来寻找能够从目标模型中诱导出特定行为的自然语言提示，但这些方法主要局限于单轮交互场景。本文研究了多轮对话背景下的行为引发问题。我们首先提出了一个分析框架，将现有方法根据其与目标模型的交互方式分为三类：仅利用先验知识的方法、利用离线交互的方法，以及从在线交互中学习的方法。随后，我们引入了在线方法的广义多轮形式化表述，统一了单轮与多轮行为引发框架。我们在自动生成多轮测试用例的任务上评估了这三类方法。通过分析查询预算（即与目标模型的交互次数）与成功率（即引发行为的输入发现率）之间的权衡关系，我们深入探究了这些方法的效率。研究发现，在三个任务中，在线方法仅需数千次查询即可实现平均45%/19%/77%的成功率，而现有多轮对话基准中的静态方法仅发现极少甚至无法找到失效案例。本研究凸显了行为引发方法在多轮对话评估中的创新应用，并表明研究社区亟需向动态基准测试方向推进。

0

相关内容

多轮对话

赋能大型语言模型多领域资源挑战

赋能大型语言模型多领域资源挑战

专知会员服务

10+阅读 · 2025年6月10日

如何将领域知识注入大模型？最新《将领域特定知识注入大语言模型》综述

如何将领域知识注入大模型？最新《将领域特定知识注入大语言模型》综述

专知会员服务

79+阅读 · 2025年2月24日

TransMLA：多头潜在注意力（MLA）即为所需

TransMLA：多头潜在注意力（MLA）即为所需

专知会员服务

23+阅读 · 2025年2月13日

迈向大语言模型偏好学习的统一视角综述

迈向大语言模型偏好学习的统一视角综述

专知会员服务

24+阅读 · 2024年9月7日

多模态大语言模型研究进展！

多模态大语言模型研究进展！

专知会员服务

42+阅读 · 2024年7月15日

基于LLM的多轮对话系统的最新进展综述

基于LLM的多轮对话系统的最新进展综述

专知会员服务

58+阅读 · 2024年3月7日

西工大等最新《大型语言模型机器人技术》综述，详述多模态 GPT-4V 机器人技术

西工大等最新《大型语言模型机器人技术》综述，详述多模态 GPT-4V 机器人技术

专知会员服务

78+阅读 · 2024年1月10日

中科大腾讯最新《多模态大型语言模型》综述，详述多模态指令微调、上下文学习、思维链和辅助视觉推理技术

中科大腾讯最新《多模态大型语言模型》综述，详述多模态指令微调、上下文学习、思维链和辅助视觉推理技术

专知会员服务

105+阅读 · 2023年6月27日

现在大火的“In-context Learning”是什么？北大等最新《语境学习ICL》综述论文，详述ICL进展、挑战和方向

现在大火的“In-context Learning”是什么？北大等最新《语境学习ICL》综述论文，详述ICL进展、挑战和方向

专知会员服务

41+阅读 · 2023年1月3日

上海交大最新《多轮对话理解》综述论文，20页pdf

上海交大最新《多轮对话理解》综述论文，20页pdf

专知会员服务

31+阅读 · 2021年10月12日

大规模跨领域中文任务导向多轮对话数据集及模型CrossWOZ

大规模跨领域中文任务导向多轮对话数据集及模型CrossWOZ

AINLP

10+阅读 · 2020年4月16日

【AAAI2020论文】多轮对话系统中的历史自适应知识融合机制, 中科院信工所孙雅静等

【AAAI2020论文】多轮对话系统中的历史自适应知识融合机制, 中科院信工所孙雅静等

专知

30+阅读 · 2019年11月24日

对话系统近期进展

对话系统近期进展

专知

37+阅读 · 2019年3月23日

深入理解BERT Transformer ，不仅仅是注意力机制

深入理解BERT Transformer ，不仅仅是注意力机制

大数据文摘

22+阅读 · 2019年3月19日

BAM！利用知识蒸馏和多任务学习构建的通用语言模型

BAM！利用知识蒸馏和多任务学习构建的通用语言模型

机器之心

15+阅读 · 2019年3月18日

【小夕精选】多轮对话之对话管理(Dialog Management)

【小夕精选】多轮对话之对话管理(Dialog Management)

夕小瑶的卖萌屋

27+阅读 · 2018年10月14日

知识在检索式对话系统的应用

知识在检索式对话系统的应用

微信AI

32+阅读 · 2018年9月20日

【论文笔记】对话模型新方法，条件DialogWAE生成多模态回答

【论文笔记】对话模型新方法，条件DialogWAE生成多模态回答

专知

15+阅读 · 2018年6月11日

多轮对话之对话管理：Dialog Management

多轮对话之对话管理：Dialog Management

PaperWeekly

18+阅读 · 2018年1月15日

赛尔原创 | 教聊天机器人进行多轮对话

赛尔原创 | 教聊天机器人进行多轮对话

哈工大SCIR

18+阅读 · 2017年9月18日

基于时空模式的复杂行为识别方法研究

国家自然科学基金

2+阅读 · 2017年12月31日

基于因子分析的会话语音说话人识别研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于人机交互的数据驱动式人群行为建模与仿真研究

国家自然科学基金

4+阅读 · 2015年12月31日

复杂决策环境下面向共识的群体评价模型与方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

主被动视角联合的细粒度行为识别

国家自然科学基金

1+阅读 · 2015年12月31日

随机环境下多个体系统集体行为分析、调控与优化

国家自然科学基金

0+阅读 · 2015年12月31日

基于犹豫模糊语言信息的定性决策理论与方法

国家自然科学基金

2+阅读 · 2015年12月31日

人类双向选择行为的统计特征分析与预测方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

考虑共谋行为的多属性采购拍卖理论与优化方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

多语言大数据环境下的复杂网络行为分析、预测和干预

国家自然科学基金

4+阅读 · 2014年12月31日

Quantifying Risks in Multi-turn Conversation with Large Language Models

Arxiv

0+阅读 · 2月4日

Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models

Arxiv

0+阅读 · 2月2日

When LLMs Play the Telephone Game: Cultural Attractors as Conceptual Tools to Evaluate LLMs in Multi-turn Settings

Arxiv

0+阅读 · 1月29日

Multimodal Conversation Structure Understanding

Arxiv

0+阅读 · 1月28日

GameTalk: Training LLMs for Strategic Conversation

Arxiv

0+阅读 · 1月22日

What Can We Actually Steer? A Multi-Behavior Study of Activation Control

Arxiv

0+阅读 · 1月11日

ACR: Adaptive Context Refactoring via Context Refactoring Operators for Multi-Turn Dialogue

Arxiv

0+阅读 · 1月9日

Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models

Arxiv

0+阅读 · 1月9日

Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey

Arxiv

0+阅读 · 1月5日

Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges

Arxiv

0+阅读 · 1月5日

VIP会员

文章信息

相关主题

相关VIP内容

赋能大型语言模型多领域资源挑战

赋能大型语言模型多领域资源挑战

专知会员服务

10+阅读 · 2025年6月10日

如何将领域知识注入大模型？最新《将领域特定知识注入大语言模型》综述

如何将领域知识注入大模型？最新《将领域特定知识注入大语言模型》综述

专知会员服务

79+阅读 · 2025年2月24日

TransMLA：多头潜在注意力（MLA）即为所需

TransMLA：多头潜在注意力（MLA）即为所需

专知会员服务

23+阅读 · 2025年2月13日

迈向大语言模型偏好学习的统一视角综述

迈向大语言模型偏好学习的统一视角综述

专知会员服务

24+阅读 · 2024年9月7日

多模态大语言模型研究进展！

多模态大语言模型研究进展！

专知会员服务

42+阅读 · 2024年7月15日

基于LLM的多轮对话系统的最新进展综述

基于LLM的多轮对话系统的最新进展综述

专知会员服务

58+阅读 · 2024年3月7日

西工大等最新《大型语言模型机器人技术》综述，详述多模态 GPT-4V 机器人技术

西工大等最新《大型语言模型机器人技术》综述，详述多模态 GPT-4V 机器人技术

专知会员服务

78+阅读 · 2024年1月10日

中科大腾讯最新《多模态大型语言模型》综述，详述多模态指令微调、上下文学习、思维链和辅助视觉推理技术

中科大腾讯最新《多模态大型语言模型》综述，详述多模态指令微调、上下文学习、思维链和辅助视觉推理技术

专知会员服务

105+阅读 · 2023年6月27日

现在大火的“In-context Learning”是什么？北大等最新《语境学习ICL》综述论文，详述ICL进展、挑战和方向

现在大火的“In-context Learning”是什么？北大等最新《语境学习ICL》综述论文，详述ICL进展、挑战和方向

专知会员服务

41+阅读 · 2023年1月3日

上海交大最新《多轮对话理解》综述论文，20页pdf

上海交大最新《多轮对话理解》综述论文，20页pdf

专知会员服务

31+阅读 · 2021年10月12日

热门VIP内容

开通专知VIP会员享更多权益服务

智能体记忆深度剖析：评价指标与系统局限性的分类体系及实证分析

《可信人工智能赋能系统的支柱》

【CMU博士论文】可靠轨迹预测的分层基石：数据、评估与方法

人工智能赋能边缘与自主系统：美陆军现代化进程聚焦威胁探测与战术边缘情报

相关资讯

大规模跨领域中文任务导向多轮对话数据集及模型CrossWOZ

大规模跨领域中文任务导向多轮对话数据集及模型CrossWOZ

AINLP

10+阅读 · 2020年4月16日

【AAAI2020论文】多轮对话系统中的历史自适应知识融合机制, 中科院信工所孙雅静等

【AAAI2020论文】多轮对话系统中的历史自适应知识融合机制, 中科院信工所孙雅静等

专知

30+阅读 · 2019年11月24日

对话系统近期进展

对话系统近期进展

专知

37+阅读 · 2019年3月23日

深入理解BERT Transformer ，不仅仅是注意力机制

深入理解BERT Transformer ，不仅仅是注意力机制

大数据文摘

22+阅读 · 2019年3月19日

BAM！利用知识蒸馏和多任务学习构建的通用语言模型

BAM！利用知识蒸馏和多任务学习构建的通用语言模型

机器之心

15+阅读 · 2019年3月18日

【小夕精选】多轮对话之对话管理(Dialog Management)

【小夕精选】多轮对话之对话管理(Dialog Management)

夕小瑶的卖萌屋

27+阅读 · 2018年10月14日

知识在检索式对话系统的应用

知识在检索式对话系统的应用

微信AI

32+阅读 · 2018年9月20日

【论文笔记】对话模型新方法，条件DialogWAE生成多模态回答

【论文笔记】对话模型新方法，条件DialogWAE生成多模态回答

专知

15+阅读 · 2018年6月11日

多轮对话之对话管理：Dialog Management

多轮对话之对话管理：Dialog Management

PaperWeekly

18+阅读 · 2018年1月15日

赛尔原创 | 教聊天机器人进行多轮对话

赛尔原创 | 教聊天机器人进行多轮对话

哈工大SCIR

18+阅读 · 2017年9月18日

相关论文

Quantifying Risks in Multi-turn Conversation with Large Language Models

Arxiv

0+阅读 · 2月4日

Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models

Arxiv

0+阅读 · 2月2日

When LLMs Play the Telephone Game: Cultural Attractors as Conceptual Tools to Evaluate LLMs in Multi-turn Settings

Arxiv

0+阅读 · 1月29日

Multimodal Conversation Structure Understanding

Arxiv

0+阅读 · 1月28日

GameTalk: Training LLMs for Strategic Conversation

Arxiv

0+阅读 · 1月22日

What Can We Actually Steer? A Multi-Behavior Study of Activation Control

Arxiv

0+阅读 · 1月11日

ACR: Adaptive Context Refactoring via Context Refactoring Operators for Multi-Turn Dialogue

Arxiv

0+阅读 · 1月9日

Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models

Arxiv

0+阅读 · 1月9日

Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey

Arxiv

0+阅读 · 1月5日

Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges

Arxiv

0+阅读 · 1月5日

相关基金

基于时空模式的复杂行为识别方法研究

国家自然科学基金

2+阅读 · 2017年12月31日

基于因子分析的会话语音说话人识别研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于人机交互的数据驱动式人群行为建模与仿真研究

国家自然科学基金

4+阅读 · 2015年12月31日

复杂决策环境下面向共识的群体评价模型与方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

主被动视角联合的细粒度行为识别

国家自然科学基金

1+阅读 · 2015年12月31日

随机环境下多个体系统集体行为分析、调控与优化

国家自然科学基金

0+阅读 · 2015年12月31日

基于犹豫模糊语言信息的定性决策理论与方法

国家自然科学基金

2+阅读 · 2015年12月31日

人类双向选择行为的统计特征分析与预测方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

考虑共谋行为的多属性采购拍卖理论与优化方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

多语言大数据环境下的复杂网络行为分析、预测和干预

国家自然科学基金

4+阅读 · 2014年12月31日

微信扫码咨询专知VIP会员