Automated prompt optimization (APO) aims to improve large language model performance by refining prompt instructions. However, existing methods are largely constrained by fixed prompt templates, limited search spaces, or single-sided optimization that treats user questions as immutable inputs. In practice, question formulation and prompt design are inherently interdependent: clearer question structures facilitate focused reasoning and task understanding, while effective prompts reveal better ways to organize and restate queries. Ignoring this coupling fundamentally limits the effectiveness and adaptability of current APO approaches. We propose a unified multi-agent system (Helix) that jointly optimizes question reformulation and prompt instructions through a structured three-stage co-evolutionary framework. Helix integrates (1) planner-guided decomposition that breaks optimization into coupled question-prompt objectives, (2) dual-track co-evolution where specialized agents iteratively refine and critique each other to produce complementary improvements, and (3) strategy-driven question generation that instantiates high-quality reformulations for robust inference. Extensive experiments on 12 benchmarks against 6 strong baselines demonstrate the effectiveness of Helix, achieving up to 3.95% performance improvements across tasks with favorable optimization efficiency.
翻译:摘要:自动化提示优化(APO)旨在通过改进提示指令来提升大语言模型性能。然而,现有方法大多受限于固定提示模板、有限搜索空间或单一视角优化(即将用户问题视为不可变输入)。在实践中,问题表述与提示设计本质上是相互依存的:更清晰的问题结构有助于聚焦推理与任务理解,而有效的提示则揭示了组织与重述查询的更优方式。忽视这种耦合关系从根本上限制了当前APO方法的有效性与适应性。我们提出一种统一的多智能体系统(Helix),通过结构化三阶段协同进化框架,联合优化问题重构与提示指令。Helix集成了:(1)规划器引导的解耦机制,将优化目标拆解为耦合的问题-提示目标;(2)双轨协同进化机制,由专业化智能体迭代改进与相互批判以产生互补性优化;(3)策略驱动的问题生成机制,为鲁棒推理实例化高质量重构。在12个基准测试中与6个强基线的广泛实验表明,Helix的有效性显著,在跨任务场景中实现了高达3.95%的性能提升,且具备优异的优化效率。