Sequential experimental design under expensive, gradient-free objectives is a central challenge in computational statistics: evaluation budgets are tightly constrained and information must be extracted efficiently from each observation. We propose \textbf{ALMAB-DC}, a GP-based sequential design framework combining active learning, multi-armed bandits (MAB), and distributed asynchronous computing for expensive black-box experimentation. A Gaussian process surrogate with uncertainty-aware acquisition identifies informative query points; a UCB or Thompson-sampling bandit controller allocates evaluations across parallel workers; and an asynchronous scheduler handles heterogeneous runtimes. We present cumulative regret bounds for the bandit components and characterize parallel scalability via Amdahl's Law. We validate ALMAB-DC on five benchmarks. On the two statistical experimental-design tasks, ALMAB-DC achieves lower simple regret than Equal Spacing, Random, and D-optimal designs in dose--response optimization, and in adaptive spatial field estimation matches the Greedy Max-Variance benchmark while outperforming Latin Hypercube Sampling; at $K=4$ the distributed setting reaches target performance in one-quarter of sequential wall-clock rounds. On three ML/engineering tasks (CIFAR-10 HPO, CFD drag minimization, MuJoCo RL), ALMAB-DC achieves 93.4\% CIFAR-10 accuracy (outperforming BOHB by 1.7\,pp and Optuna by 1.1\,pp), reduces airfoil drag to $C_D = 0.059$ (36.9\% below Grid Search), and improves RL return by 50\% over Grid Search. All advantages over non-ALMAB baselines are statistically significant under Bonferroni-corrected Mann--Whitney $U$ tests. Distributed execution achieves $7.5\times$ speedup at $K = 16$ agents, consistent with Amdahl's Law.


翻译:在昂贵且无梯度代价函数下的序贯实验设计是计算统计学中的核心挑战:评估预算严格受限,且必须从每次观测中高效提取信息。本文提出\textbf{ALMAB-DC}——一种融合主动学习、多臂老虎机(MAB)与分布式异步计算的高斯过程(GP)序贯设计框架,适用于昂贵黑箱实验。该框架通过具有不确定性感知采集函数的高斯过程代理模型识别信息量最大的查询点;利用上置信界(UCB)或汤普森采样策略的虎机控制器在并行工作节点间分配评估任务;并采用异步调度器处理异构运行时。我们推导了虎机组件的累积遗憾界,并依据阿姆达尔定律刻画了并行可扩展性。在五个基准任务上验证了ALMAB-DC的性能:两项统计实验设计任务中,ALMAB-DC在剂量-响应优化中的简单遗憾低于等间距、随机与D最优设计;在自适应空间场估计中,其性能与贪婪最大方差基准相当且优于拉丁超立方采样;当$K=4$时,分布式设置达到目标性能所需的序贯墙钟时间仅为四分之一。三项机器学习/工程任务(CIFAR-10超参数优化、CFD阻力最小化、MuJoCo强化学习)中,ALMAB-DC在CIFAR-10上达到93.4%准确率(超BOHB 1.7个百分点与Optuna 1.1个百分点),将翼型阻力降至$C_D = 0.059$(比网格搜索低36.9%),并在强化学习回报上较网格搜索提升50%。基于Bonferroni校正的Mann-Whitney $U$检验表明,所有相对于非ALMAB基线的优势均具有统计显著性。分布式执行在$K=16$个代理时达到$7.5\times$加速比,符合阿姆达尔定律预测。

0
下载
关闭预览

相关内容

DC:Distributed Computing。 Explanation:分布式计算。 Publisher:Springer。 SIT:http://dblp.uni-trier.de/db/journals/dc/
Meta-Transformer:多模态学习的统一框架
专知会员服务
59+阅读 · 2023年7月21日
《分布式多智能体强化学习的编码》加州大学等
专知会员服务
57+阅读 · 2022年11月2日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
类脑计算的前沿论文,看我们推荐的这7篇
人工智能前沿讲习班
21+阅读 · 2019年1月7日
Spark机器学习:矩阵及推荐算法
LibRec智能推荐
16+阅读 · 2017年8月3日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
论文 | 分子性质预测中的闭环自动研究与泛化认证
论文 | OmniScientist:全模态全学科AI科学家
专知会员服务
0+阅读 · 今天14:40
无人机已改变战场,但并未解决指挥问题
专知会员服务
5+阅读 · 8月14日
驱动军事决策变革的顶尖人工智能指挥系统
专知会员服务
11+阅读 · 8月11日
非对称防御中的自组织临界性:俄乌战争
专知会员服务
10+阅读 · 8月10日
相关VIP内容
Meta-Transformer:多模态学习的统一框架
专知会员服务
59+阅读 · 2023年7月21日
《分布式多智能体强化学习的编码》加州大学等
专知会员服务
57+阅读 · 2022年11月2日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员