Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. Current methods either rely on disjoint pipelines, which fail to capture fine-grained interactions, or require structured inputs and external text rewriting, which limits the flexibility of free-form text prompts. In this paper, we introduce a new task: Free-Form-Text-Prompt-to-Unified-Audio generation, which aims to directly synthesize unified audio containing speech, sound, and their composites from unconstrained natural language. To address this task, we propose PlanAudio, a unified, autoregressive LLM-based framework. First, it simplifies the model architecture by leveraging intrinsic LLM reasoning capability instead of traditional text encoders. Second, it introduces a semantic latent chain-of-thought mechanism, an implicit planning mechanism that bridges high-level semantic understanding and low-level acoustic synthesis. Furthermore, we create PlanAudio-Bench, a specialized benchmark for evaluating composite audio scenarios. We perform evaluations in the scenarios of speech, sound, and their composites. The results demonstrate that PlanAudio generally outperforms the existing pipeline and unified baselines, while staying competitive with models designed for a single scenario. Our analysis further reveals the superiority of semantic latent CoT over other CoT mechanisms and highlights the importance of continuous multi-scenario training curricula.
翻译:音频生成已取得显著进展,但将语音与声音自然合成为统一音频仍是一项挑战。当前方法要么依赖无法捕捉细粒度交互的解耦流水线,要么需要结构化输入和外部文本改写,这限制了自由文本提示的灵活性。本文提出一项新任务:自由文本提示到统一音频生成,旨在直接从无约束自然语言合成包含语音、声音及其复合物的统一音频。为应对此任务,我们提出PlanAudio——一种统一的、基于自回归大语言模型(LLM)的框架。首先,该框架利用LLM内在推理能力替代传统文本编码器,简化模型架构。其次,引入语义潜在思维链(semantic latent chain-of-thought)机制,这是一种隐式规划机制,桥接了高层语义理解与低层声学合成。此外,我们构建了PlanAudio-Bench,一个专门评估复合音频场景的基准数据集。我们在语音、声音及其复合场景下进行评估。结果表明,PlanAudio普遍优于现有流水线及统一基线方法,同时与专为单场景设计的模型保持竞争力。进一步分析揭示了语义潜在CoT相较于其他CoT机制的优越性,并强调了连续多场景训练课程的重要性。