Current Large Language Models (LLMs) are gradually exploited in practically valuable agentic workflows such as Deep Research, E-commerce recommendation, and job recruitment. In these applications, LLMs need to select some optimal solutions from massive candidates, which we term as \textit{LLM-as-a-Recommender} paradigm. However, the reliability of using LLM agents for recommendations is underexplored. In this work, we introduce a \textbf{Bias} \textbf{Rec}ommendation \textbf{Bench}mark (\textbf{BiasRecBench}) to highlight the critical vulnerability of such agents to biases in high-value real-world tasks. The benchmark includes three practical domains: paper review, e-commerce, and job recruitment. We construct a \textsc{Bias Synthesis Pipeline with Calibrated Quality Margins} that 1) synthesizes evaluation data by controlling the quality gap between optimal and sub-optimal options to provide a calibrated testbed to elicit the vulnerability to biases; 2) injects contextual biases that are logical and suitable for option contexts. Extensive experiments on both SOTA (Gemini-{2.5,3}-pro, GPT-4o, DeepSeek-R1) and small-scale LLMs reveal that agents frequently succumb to injected biases despite having sufficient reasoning capabilities to identify the ground truth. These findings expose a significant reliability bottleneck in current agentic workflows, calling for specialized alignment strategies for LLM-as-a-Recommender. The complete code and evaluation datasets will be made publicly available shortly.


翻译:当前大型语言模型(LLM)正逐步被应用于深度研究、电商推荐和人才招聘等实际有价值的智能体工作流中。在此类应用中,LLM需从海量候选中筛选最优解,我们称之为“LLM即推荐代理”(LLM-as-a-Recommender)范式。然而,LLM代理用于推荐的可靠性尚未得到充分探索。本文提出**偏差推荐基准**(**BiasRecBench**),旨在揭示此类代理在高价值真实世界任务中对偏见的严重脆弱性。该基准涵盖三大实践领域:论文评审、电子商务和人才招聘。我们构建了一个**具有校准质量边际的偏差合成流水线**(Bias Synthesis Pipeline with Calibrated Quality Margins):1)通过控制最优与次优选项间的质量差距来合成评估数据,提供校准测试平台以诱发对偏见的脆弱性;2)注入逻辑合理且契合选项上下文的语境偏差。针对当前最优(SOTA)模型(Gemini-{2.5,3}-pro、GPT-4o、DeepSeek-R1)及小规模LLM的广泛实验表明,尽管智能体具备识别真实答案的充分推理能力,却频繁屈服于注入的偏见。这些发现暴露了当前智能体工作流中显著的可靠性瓶颈,亟需针对LLM即推荐代理范式开发专门的鲁棒性对齐策略。完整代码与评估数据集将于近期公开。

0
下载
关闭预览

相关内容

LLM/智能体作为数据分析师:综述
专知会员服务
39+阅读 · 2025年9月30日
关于大语言模型驱动的推荐系统智能体的综述
专知会员服务
31+阅读 · 2025年2月17日
【ICLR2024】能检测到LLM产生的错误信息吗?
专知会员服务
25+阅读 · 2024年1月23日
最全推荐系统Embedding召回算法总结
凡人机器学习
30+阅读 · 2020年7月5日
我是怎么走上推荐系统这条(不归)路的……
全球人工智能
11+阅读 · 2019年4月9日
深度 | 推荐系统评估
AI100
24+阅读 · 2019年3月16日
推荐系统
炼数成金订阅号
28+阅读 · 2019年1月17日
推荐系统概述
Python开发者
11+阅读 · 2018年9月27日
放弃 RNN/LSTM 吧,因为真的不好用!望周知~
人工智能头条
19+阅读 · 2018年4月24日
推荐系统杂谈
架构文摘
28+阅读 · 2017年9月15日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
《人工智能赋能的适应性多功能电磁战》
专知会员服务
7+阅读 · 9月29日
俄乌战场实验室:全面战争如何重塑现代作战
专知会员服务
5+阅读 · 9月29日
2026年美空军协会会议上的无人机系统趋势
专知会员服务
8+阅读 · 9月28日
反制无人机:乌克兰提供的五点启示
专知会员服务
12+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
10+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
9+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
13+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
17+阅读 · 9月22日
相关VIP内容
LLM/智能体作为数据分析师:综述
专知会员服务
39+阅读 · 2025年9月30日
关于大语言模型驱动的推荐系统智能体的综述
专知会员服务
31+阅读 · 2025年2月17日
【ICLR2024】能检测到LLM产生的错误信息吗?
专知会员服务
25+阅读 · 2024年1月23日
相关资讯
最全推荐系统Embedding召回算法总结
凡人机器学习
30+阅读 · 2020年7月5日
我是怎么走上推荐系统这条(不归)路的……
全球人工智能
11+阅读 · 2019年4月9日
深度 | 推荐系统评估
AI100
24+阅读 · 2019年3月16日
推荐系统
炼数成金订阅号
28+阅读 · 2019年1月17日
推荐系统概述
Python开发者
11+阅读 · 2018年9月27日
放弃 RNN/LSTM 吧,因为真的不好用!望周知~
人工智能头条
19+阅读 · 2018年4月24日
推荐系统杂谈
架构文摘
28+阅读 · 2017年9月15日
相关基金
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员