Food assistance referral requires conversational agents to translate underspecified, often noisy help-seeking dialogues into locally valid resource recommendations. We present Food4All, an agentic food-resource referral framework and benchmark grounded in 686 structured Indiana food resources. Food4All couples a food-specific search tool with 300 multi-turn evaluation tasks spanning single food needs, composite cases with access or document constraints, and five non-ideal user interaction traits: unreasonable demands, rambling responses, impatience, incomplete answers, and inconsistent information. We evaluate six Large Language Models (LLMs) on requirement grounding, resource retrieval, final referral correctness, and interaction efficiency. Although the strongest model achieves 96.33% referral accuracy, our diagnostics reveal persistent failures in grounding schedule, eligibility, intake, and document constraints, as well as failures to preserve valid retrieved resources in the final recommendation. Trait-level analysis further shows that different non-ideal behaviors stress different parts of the referral pipeline. Food4All provides a controlled testbed for studying tool-calling agents in constraint-sensitive food assistance referral under realistic user interaction challenges.
翻译:食品援助转介要求对话智能体将表述模糊且常含噪音的求助对话转化为符合当地实际情况的资源推荐。我们提出Food4All——一个基于印第安纳州686条结构化食品资源的智能体食品资源转介框架与基准。Food4All将食品专用搜索工具与300个多轮评估任务相结合,涵盖单一食品需求、附带准入或证件限制的复合案例,以及五种非理想用户交互特征:不合理要求、冗长回应、不耐烦、回答不完整和信息不一致。我们在需求锚定、资源检索、最终转介正确性和交互效率四个维度上评估了六个大语言模型。尽管最强模型达到了96.33%的转介准确率,但诊断分析揭示了其在锚定时间安排、资格条件、申请流程和证件限制方面的持续性失败,以及未能将有效检索资源保留至最终推荐的问题。特质层面分析进一步表明,不同非理想行为对转介流程的不同环节产生差异化压力。Food4All为在真实用户交互挑战下研究约束敏感型食品援助转介中的工具调用智能体提供了受控测试平台。