Large language model agents are increasingly integrated into map services. Since map services are embedded in everyday-life scenarios rather than professional task settings, users often express their needs informally, resulting in underspecified queries with many unspoken needs, namely, implicit decision factors that are critical for user satisfaction. Although clarification is an effective way to mitigate this issue, it increases user burden in daily interaction, and a capable agent should first proactively recover such factors from available information sources. However, evaluating this ability is challenging. The first challenge is to determine which implicit decision factors are suitable for evaluation. A factor is evaluable only if it affects user acceptance and can be recovered from information available to the agent before it responds. Second, user satisfaction cannot be reliably represented by a single reference answer, requiring a benchmark that converts satisfaction-relevant factors into objective and quantifiable evaluation targets. To address these challenges, we propose a restore-identify-filter framework that reconstructs complete user needs from behavior-chain evidence, identifies implicit decision factors, and retains only those supported by pre-query evidence. Building on this methodology, we construct MapSatisfyBench from large-scale, real-world anonymized user data and annotate ground truth from five dimensions and enables full-chain evaluation of satisfaction-aware map agents. Experiments show that current agents generally perform well on explicit task completion, but remain limited in satisfying implicit decision factors and proactively acquiring the evidence needed for satisfaction-aware decisions. These findings establish MapSatisfyBench as a benchmark for shifting map-agent evaluation from task completion toward satisfaction-aware spatial decision making.
翻译:大语言模型代理正日益集成到地图服务中。由于地图服务嵌入日常生活场景而非专业任务设置,用户往往以非正式方式表达需求,导致查询描述不充分,包含许多未言明的需求,即对用户满意度至关重要的隐含决策因素。尽管澄清是缓解此问题的有效方式,但会增加日常交互中用户的负担,而有能力的代理应首先主动从可用信息源中恢复这些因素。然而,评估这一能力颇具挑战。首要挑战是确定哪些隐含决策因素适合评估。一个因素仅在影响用户接受度且能从代理响应前可获取的信息中恢复时,才具备可评估性。其次,用户满意度无法由单一参考答案可靠表示,这需要一个将满意度相关因素转化为客观可量化评估目标的基准。为应对这些挑战,我们提出一个“恢复-识别-筛选”框架:从行为链证据中重构完整的用户需求,识别隐含决策因素,并仅保留那些有查询前证据支持的因素。基于此方法论,我们利用大规模真实世界匿名化用户数据构建了MapSatisfyBench,并从五个维度标注真值,实现对满意度感知地图代理的全链条评估。实验表明,当前代理在显式任务完成上普遍表现良好,但在满足隐含决策因素及主动获取满意度感知决策所需的证据方面仍存在局限。这些发现确立了MapSatisfyBench作为基准,推动地图代理评估从任务完成向满意度感知空间决策转变。