Reinforcement learning fine-tuning has proven effective for steering generative diffusion models toward desired properties in image and molecular domains. Graph diffusion models have similarly been applied to combinatorial structure generation, including neural architecture search (NAS). However, neural architectures are directed acyclic graphs (DAGs) where edge direction encodes functional semantics such as data flow-information that existing graph diffusion methods, designed for undirected structures, discard. We propose Directed Graph Policy Optimization (DGPO), which extends reinforcement learning fine-tuning of discrete graph diffusion models to DAGs via topological node ordering and positional encoding. Validated on NAS-Bench-101 and NAS-Bench-201, DGPO matches the benchmark optimum on all three NAS-Bench-201 tasks (91.61%, 73.49%, 46.77%). The central finding is that the model learns transferable structural priors: pretrained on only 7% of the search space, it generates near-oracle architectures after fine-tuning, within 0.32 percentage points of the full-data model and extrapolating 7.3 percentage points beyond its training ceiling. Bidirectional control experiments confirm genuine reward-driven steering, with inverse optimization reaching near random-chance accuracy (9.5%). These results demonstrate that reinforcement learning-steered discrete diffusion, once extended to handle directionality, provides a controllable generative framework for directed combinatorial structures.


翻译:强化学习微调已被证明能有效引导生成扩散模型在图像和分子领域生成符合预期属性的样本。图扩散模型同样被应用于组合结构生成,包括神经架构搜索(NAS)。然而,神经架构是有向无环图(DAG),其边方向编码了数据流等功能语义——现有面向无向结构设计的图扩散方法会忽略此类信息。我们提出有向图策略优化(DGPO),该方法通过拓扑节点排序和位置编码将离散图扩散模型的强化学习微调扩展至DAG。在NAS-Bench-101和NAS-Bench-201上的验证表明,DGPO在所有三个NAS-Bench-201任务中均达到基准最优值(91.61%、73.49%、46.77%)。核心发现是模型可学习可迁移的结构先验:仅使用搜索空间7%的数据预训练后,微调即可生成近乎最优架构,与全数据模型仅差0.32个百分点,并超越其训练上限7.3个百分点。双向控制实验证实了真正的奖励驱动引导,逆优化仅达到接近随机准确性(9.5%)。这些结果表明,强化学习引导的离散扩散在扩展至处理方向性后,为有向组合结构提供了可控生成框架。

0
下载
关闭预览

相关内容

图增强生成(GraphRAG)
专知会员服务
35+阅读 · 2025年1月4日
【Google AI】鲁棒图神经网络,Robust Graph Neural Networks
专知会员服务
38+阅读 · 2022年3月9日
综述| 当图神经网络遇上强化学习
图与推荐
35+阅读 · 2022年7月1日
当深度强化学习遇见图神经网络
专知
227+阅读 · 2019年10月21日
【GNN】深度学习之上,图神经网络(GNN )崛起
产业智能官
16+阅读 · 2019年8月15日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
41+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Arxiv
22+阅读 · 2023年11月2日
Arxiv
18+阅读 · 2019年3月28日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
3+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
7+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关VIP内容
图增强生成(GraphRAG)
专知会员服务
35+阅读 · 2025年1月4日
【Google AI】鲁棒图神经网络,Robust Graph Neural Networks
专知会员服务
38+阅读 · 2022年3月9日
相关基金
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
41+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员