The integration of Large Language Model (LLM) agents is transforming recommender systems from simple query-item matching towards deeply personalized and interactive recommendations. Reinforcement Learning (RL) provides an essential framework for the optimization of these agents in recommendation tasks. However, current methodologies remain limited by a reliance on single dimensional outcome-based rewards that focus exclusively on final user interactions, overlooking critical intermediate capabilities, such as instruction following and complex intent understanding. Despite the necessity for designing multi-dimensional reward, the field lacks a standardized benchmark to facilitate this development. To bridge this gap, we introduce RecRM-Bench, the largest and most comprehensive benchmark to date for agentic recommender systems. It comprises over 1 million structured entries across four core evaluation dimensions: instruction following, factual consistency, query-item relevance, and fine-grained user behavior prediction. By supporting comprehensive assessment from syntactic compliance to complex intent grounding and preference modeling, RecRM-Bench provides a foundational dataset for training sophisticated reward models. Furthermore, we propose a systematic framework for the construction of multi-dimensional reward models and the integration of a hybrid reward function, establishing a robust foundation for developing reliable and highly capable agentic recommender systems. The complete RecRM-Bench dataset is publicly available at https://huggingface.co/datasets/wwzeng/RecRM-Bench.
翻译:大型语言模型(LLM)智能体的集成正将推荐系统从简单的查询-商品匹配重构为深度个性化与交互式推荐。强化学习(RL)为优化这些智能体在推荐任务中的表现提供了核心框架。然而,当前方法仍受限于依赖仅关注最终用户交互的单维度结果奖励,忽视了指令遵循与复杂意图理解等关键中间能力。尽管设计多维度奖励势在必行,该领域仍缺乏标准化基准以支撑相关研究。为填补这一空白,我们提出RecRM-Bench——迄今为止规模最大、维度最全面的智能推荐系统基准测试。它包含超过100万条结构化数据,覆盖四个核心评估维度:指令遵循、事实一致性、查询-商品相关性以及细粒度用户行为预测。通过支持从语法合规性到复杂意图根植及偏好建模的全面评估,RecRM-Bench为训练高阶奖励模型提供了基础数据集。此外,我们提出了一套系统化框架,用于构建多维度奖励模型并集成混合奖励函数,为开发可靠且高性能的智能推荐系统奠定了坚实基础。完整RecRM-Bench数据集已公开于 https://huggingface.co/datasets/wwzeng/RecRM-Bench。