We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via Serialized Output Training (SOT) to learn turn-taking dynamics; and (2) an interleaved time anchor mechanism that not only supports fine-grained timestamp prediction but also acts as a synchronization signal between semantic understanding and speaker tracking. Compared to previous works that primarily focus on speaker-attributed ASR or implicit diarization, TagSpeech addresses the challenge of fine-grained speaker-content alignment and explicitly models "who spoke what and when" in an end-to-end manner. Experiments on AMI and AliMeeting benchmarks demonstrate that our method achieves consistent improvements in Diarization Error Rate (DER) over strong end-to-end baselines, including Qwen-Omni and Gemini, particularly in handling complex speech overlaps. Moreover, TagSpeech employs a parameter-efficient training paradigm in which the LLM backbone is frozen and only lightweight projectors are trained, resulting in strong performance with low computational cost.


翻译:本文提出TagSpeech,一种基于大语言模型(LLM)的统一框架,利用时序锚定机制实现多说话人语音识别与说话人日志的联合建模。该框架基于两个核心设计:(1)通过串行输出训练(SOT)微调的解耦语义流与说话人流,以学习话轮转换动态;(2)交错式时间锚机制,不仅支持细粒度时间戳预测,同时作为语义理解与说话人追踪之间的同步信号。相较于以往主要关注说话人归属语音识别或隐式日志的研究,TagSpeech解决了细粒度说话人-内容对齐的挑战,并以端到端方式显式建模“何人于何时说出何内容”。在AMI与AliMeeting基准测试上的实验表明,本方法在说话人日志错误率(DER)上相较于包括Qwen-Omni与Gemini在内的强端到端基线模型取得持续提升,尤其在处理复杂语音重叠场景中表现突出。此外,TagSpeech采用参数高效的训练范式,冻结LLM主干网络仅训练轻量级投影器,从而以较低计算成本实现强劲性能。

0
下载
关闭预览

相关内容

《深度学习多标签学习》最新综述
专知会员服务
48+阅读 · 2024年1月31日
AI Agent,大模型时代重要落地方向, 42页ppt
专知会员服务
292+阅读 · 2023年10月12日
【CVPR2022】语言作为查询的参考视频目标分割框架
专知会员服务
10+阅读 · 2022年4月27日
重磅发布:基于 PyTorch 的深度文本匹配工具 MatchZoo-py
中国科学院网络数据重点实验室
16+阅读 · 2019年8月26日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
VIP会员
最新内容
刚刚!Jev中文教程项目发布了
专知会员服务
0+阅读 · 10月4日
《人工智能赋能的适应性多功能电磁战》
专知会员服务
12+阅读 · 9月29日
俄乌战场实验室:全面战争如何重塑现代作战
专知会员服务
8+阅读 · 9月29日
2026年美空军协会会议上的无人机系统趋势
专知会员服务
11+阅读 · 9月28日
反制无人机:乌克兰提供的五点启示
专知会员服务
17+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
11+阅读 · 9月23日
相关VIP内容
《深度学习多标签学习》最新综述
专知会员服务
48+阅读 · 2024年1月31日
AI Agent,大模型时代重要落地方向, 42页ppt
专知会员服务
292+阅读 · 2023年10月12日
【CVPR2022】语言作为查询的参考视频目标分割框架
专知会员服务
10+阅读 · 2022年4月27日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员