We study zeroth-order optimization where solutions must minimize a cost $d(s)$ while maintaining high probability under a complex generative prior $L(s)$ (e.g., a parameterized model). This reduces to sampling from a target distribution proportional to $L(s) e^{-T \cdot d(s)}$. Since classical model-based optimization (MBO) lacks finite-sample guarantees for expressive approximate learners, we introduce "coarse learnability", a flexible statistical assumption requiring only that a learned model covers the target's probability mass within a polynomial factor. Leveraging this assumption, we design an iterative MBO algorithm called \alift with a sample correction step that provably approximates the target using only a polynomial number of samples. We apply this framework to globally optimizing non-convex objectives bounded by a quadratic envelope in $R^n$, where we show this assumption is naturally satisfied for a family of "optimistic" posterior distributions. To reach global $\varepsilon$-optimality, this implies a sample complexity of $\widetilde{O}(\log 1/\varepsilon)$, a rate characteristic of optimistic space-partitioning methods. We further justify coarse learnability as an assumption for generative priors theoretically, proving that in simple settings, parametric maximum likelihood estimation and over-smoothed kernel density estimators naturally satisfy it. Finally, one motivation for our framework comes from inference-time alignment. Though our primary contribution pertains to the theoretical foundations of MBO, we provide qualitative evidence that, in simple settings, even primitive LLMs can shift their distributions toward lower-cost regions when fine-tuned with zeroth-order feedback.


翻译:我们研究零阶优化问题,其中解必须在最小化代价函数$d(s)$的同时,以高概率保持与复杂生成先验$L(s)$(如参数化模型)的一致性。这等价于从与$L(s) e^{-T \cdot d(s)}$成正比的目標分布中采样。由于经典基于模型的优化(MBO)对表达能力强的近似学习器缺乏有限样本保证,我们引入了“粗可学习性”——一种灵活的统计假设,仅要求学习模型在多项式因子范围内覆盖目标分布的概率质量。基于这一假设,我们设计了名为\alift的迭代式MBO算法,该算法通过样本校正步骤,仅需多项式数量的样本即可证明性地逼近目标分布。我们将该框架应用于全局优化具有二次包络界的非凸目标函数(在$\mathbb{R}^n$空间中),并证明在此类问题中,对于一族“乐观”后验分布,该假设自然成立。为实现全局$\varepsilon$-最优性,该方法的样本复杂度为$\widetilde{O}(\log 1/\varepsilon)$,这一速率是乐观空间划分方法的典型特征。我们进一步从理论上论证了粗可学习性作为生成先验假设的合理性,证明在简单场景中,参数化极大似然估计和过度平滑的核密度估计自然满足该假设。最后,我们框架的动机之一来自推理时对齐。尽管我们的主要贡献在于MBO的理论基础,但我们提供的定性证据表明:在简单设定下,即使原始的大型语言模型(LLM)通过零阶反馈微调后,也能将其分布向低代价区域迁移。

0
下载
关闭预览

相关内容

【阿姆斯特丹博士论文】带约束学习的优化算法
专知会员服务
21+阅读 · 2025年4月4日
多样化偏好优化
专知会员服务
12+阅读 · 2025年2月3日
机器学习中的最优化算法总结
人工智能前沿讲习班
22+阅读 · 2019年3月22日
入门 | 深度学习模型的简单优化技巧
机器之心
10+阅读 · 2018年6月10日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月12日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员