Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant to long-context reasoning pose well-known challenges for automatic evaluation. We propose a synthetic data generation pipeline designed to serve both as a training resource and as a controlled evaluation environment, and instantiate it for first-visit doctor-patient conversations with SOAP note generation as the task. The pipeline has three stages, persona-driven dialogue generation, multi-speaker audio synthesis with overlap/pause modeling, room acoustics, and sound events, and LLM-based reference SOAP note production, built entirely on open-weight models. We release 8,800 synthetic conversations with 1.3k hours of corresponding audio and reference notes. Evaluating current open-weight systems, we find that cascaded approaches still substantially outperform end-to-end models.
翻译:长上下文音频推理在训练数据和评估两方面均存在不足。现有基准主要针对短上下文任务,而与长上下文推理最相关的开放式生成任务在自动评估方面面临公认的挑战。我们提出了一种合成数据生成流水线,旨在同时作为训练资源和受控评估环境,并以初诊医患对话为实例,以SOAP笔记生成为任务。该流水线包含三个阶段:基于人物设定的对话生成、具备重叠/停顿建模、房间声学及声音事件的多说话人音频合成,以及基于大语言模型的参考SOAP笔记制作,全部基于开源权重模型构建。我们发布了8,800段合成对话,包含1,300小时对应音频及参考笔记。对现有开源系统评估发现,级联方法仍显著优于端到端模型。