We introduce the first Natural Office Talkers in Settings of Far-field Audio Recordings (``NOTSOFAR-1'') Challenge alongside datasets and baseline system. The challenge focuses on distant speaker diarization and automatic speech recognition (DASR) in far-field meeting scenarios, with single-channel and known-geometry multi-channel tracks, and serves as a launch platform for two new datasets: First, a benchmarking dataset of 315 meetings, averaging 6 minutes each, capturing a broad spectrum of real-world acoustic conditions and conversational dynamics. It is recorded across 30 conference rooms, featuring 4-8 attendees and a total of 35 unique speakers. Second, a 1000-hour simulated training dataset, synthesized with enhanced authenticity for real-world generalization, incorporating 15,000 real acoustic transfer functions. The tasks focus on single-device DASR, where multi-channel devices always share the same known geometry. This is aligned with common setups in actual conference rooms, and avoids technical complexities associated with multi-device tasks. It also allows for the development of geometry-specific solutions. The NOTSOFAR-1 Challenge aims to advance research in the field of distant conversational speech recognition, providing key resources to unlock the potential of data-driven methods, which we believe are currently constrained by the absence of comprehensive high-quality training and benchmarking datasets.
翻译:我们介绍了首个自然办公场景远场音频录制环境中的语音交流挑战赛(简称“NOTSOFAR-1”),同时发布相关数据集与基线系统。该挑战聚焦于远场会议场景下的远场说话人日志与自动语音识别(DASR),涵盖单通道与已知几何结构的的多通道赛道,并作为两个新数据集的发布平台:其一为包含315场会议(平均时长6分钟)的基准测试数据集,覆盖广泛真实声学环境与对话动态,这些会议在30间会议室中录制,参与者4-8人,总计包含35位不同说话人;其二为1000小时模拟训练数据集,通过增强真实性策略实现真实场景泛化,并融入15000个真实声学传递函数。任务专注于单设备DASR,其中多通道设备始终共享相同的已知几何结构,这一设定与实际会议室常见配置一致,避免了多设备任务的技术复杂度,同时支持开发针对特定几何结构的解决方案。NOTSOFAR-1挑战赛旨在推动远场会话语音识别领域的研究,通过提供关键资源以释放数据驱动方法的潜力——我们认为该方法目前正受限于高质量训练与基准测试数据集的缺失。