Forecasting real-world events requires language-model agents to reason under uncertainty from incomplete, time-bounded information. Yet evaluating whether agents genuinely forecast requires more than final-answer accuracy: a model may be correct by recalling memorized training facts, citing fabricated evidence, or producing an unsupported causal story. We present WorldReasoner, an evaluation framework for temporally valid event forecasting. Each task gives an agent a resolved forecasting question, a simulated forecast date, and access only to evidence available before that date; after resolution, the framework scores the submitted probability, cited evidence, and optional causal event graph. WorldReasoner reports three complementary axes: outcome quality against resolved answers, evidence quality over cited sources, and reasoning quality against post-resolution hindsight graphs. The benchmark is built by an agentic construction pipeline that generates forecasting questions, collects time-stamped evidence, and builds hindsight reference graphs at scale, yielding 345 resolved tasks derived from 14,141 articles with graphs covering 8,087 extracted events. Across six controlled agent settings, temporally valid retrieval is the strongest driver of outcome accuracy; causal graph construction improves key-event recovery; and correct graph-enabled forecasts are more strongly grounded in key events and relevant sources, yet agents still struggle to convert grounded evidence into calibrated probabilities.


翻译:预测现实世界事件要求语言模型智能体基于不完整且有时间约束的信息,在不确定性下进行推理。然而,评估智能体是否真正进行预测,仅凭最终答案的准确性远远不够:模型可能因回忆记忆中的训练事实、引用捏造的证据或提出无据可依的因果叙事而正确。我们提出WorldReasoner,一个用于时间有效事件预测的评估框架。每个任务为智能体提供一个已解决预测问题、一个模拟预测日期,并仅允许访问该日期之前的可用证据;在事件解决后,该框架对提交的概率、引用的证据以及可选的因果事件图进行评分。WorldReasoner报告三个互补维度:针对已解决答案的结果质量、针对引用来源的证据质量,以及针对事后回溯图的推理质量。该基准通过一个代理式构建管道生成,该管道大规模生成预测问题、收集带时间戳的证据并构建事后参考图,最终从14,141篇文章中提取345个已解决任务,其因果图覆盖8,087个抽取事件。在六个受控智能体设置下,时间有效检索是结果准确性的最强驱动因素;因果图构建可提高关键事件恢复能力;图启用的正确预测更强地扎根于关键事件和相关来源,但智能体仍难以将扎实的证据转化为校准的概率。

0
下载
关闭预览

相关内容

大语言模型的智能体化推理
专知会员服务
36+阅读 · 1月21日
具身智能中的世界模型:全面综述
专知会员服务
55+阅读 · 2025年10月21日
走向通用人工智能之路,世界模型为何不可或缺?
专知会员服务
21+阅读 · 2025年7月1日
理解世界还是预测未来?世界模型的综合综述
专知会员服务
79+阅读 · 2024年11月26日
「大型语言模型推理」综述
专知会员服务
96+阅读 · 2022年12月24日
「知识增强预训练语言模型」最新研究综述
专知
18+阅读 · 2022年11月18日
自然语言处理中的语言模型预训练方法
PaperWeekly
14+阅读 · 2018年10月21日
回归预测&时间序列预测
GBASE数据工程部数据团队
44+阅读 · 2017年5月17日
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Arxiv
3+阅读 · 6月11日
Arxiv
14+阅读 · 2023年8月7日
VIP会员
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
4+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
5+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
4+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
7+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
10+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
5+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
10+阅读 · 9月21日
相关VIP内容
大语言模型的智能体化推理
专知会员服务
36+阅读 · 1月21日
具身智能中的世界模型:全面综述
专知会员服务
55+阅读 · 2025年10月21日
走向通用人工智能之路,世界模型为何不可或缺?
专知会员服务
21+阅读 · 2025年7月1日
理解世界还是预测未来?世界模型的综合综述
专知会员服务
79+阅读 · 2024年11月26日
「大型语言模型推理」综述
专知会员服务
96+阅读 · 2022年12月24日
相关基金
国家自然科学基金
10+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员