AlphaZero is normally evaluated as one agent: a policy-value network fused with Monte Carlo tree search. That fusion hides a causal question. When self-play search is given a useful prior, does the network absorb the induced behavior, or does the behavior stay rented from search at test time? We answer with Cross-Phase Prior Intervention (CPI), which switches a root-level tactical prior on and off independently during training and during evaluation, separating the prior's online effect from the learned residual it leaves in the weights. The endpoint is deliberately narrow: how often a network discharges a forced defensive obligation when no search-time guidance is available. On a sealed one-shot final test in 9x9 Gomoku and 19x19 Go, deleting the prior still leaves a large residual response rises from 13.8% to 26.3% in Gomoku and from 0.6% to 33.8% in Go-and soft reweighting teaches as well as hard action restriction, so pruning legal actions is not the mechanism. The same cross bounds the claim: re-enabling the prior restores nearly 100% response, leaving dependence gaps of 73.7 and 65.8 points. A latched-position evaluation localizes the residual to trained geometry-absent at the shared initialization, emerging over training, worth +17.3 points on in-distribution defenses but only +1.3 on structurally novel ones. Search is therefore best read as a training-time behavioral curriculum whose lessons are real, partial, and geometry-bound, and online competence and internalized competence are different estimands that a diagonal ablation cannot tell apart.


翻译:暂无翻译

0
下载
关闭预览

相关内容

BES:让语言模型通过双向进化搜索自我改进
专知会员服务
9+阅读 · 5月30日
AlphaMosaic:人工智能赋能的作战管理系统
专知会员服务
46+阅读 · 2025年8月19日
AlphaFold教程与最新蛋白质结构预测进展,附视频与Slides
专知会员服务
29+阅读 · 2022年6月16日
AlphaZero原理与启示
专知会员服务
33+阅读 · 2020年8月23日
搜索query意图识别的演进
DataFunTalk
13+阅读 · 2020年11月15日
论文浅尝 | 基于知识图谱子图匹配以回答自然语言问题
开放知识图谱
26+阅读 · 2018年6月26日
Reinforcement Learning: An Introduction 2018第二版 500页
CreateAMind
14+阅读 · 2018年4月27日
论文 | YOLO(You Only Look Once)目标检测
七月在线实验室
14+阅读 · 2017年12月12日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Arxiv
0+阅读 · 8月12日
Arxiv
14+阅读 · 2021年11月27日
VIP会员
最新内容
《最强大的军事网状网络》
专知会员服务
4+阅读 · 9月7日
《预测陆军征兵任务分配》110页
专知会员服务
4+阅读 · 9月7日
分层反无人机系统发展新趋势
专知会员服务
11+阅读 · 9月3日
何为协作武器?
专知会员服务
11+阅读 · 9月1日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员