Target speaker extraction (TSE) extracts the target speaker's voice from overlapping speech given a reference utterance. Existing masking-based approaches are lightweight and effective but suffer from an inability to synthesize missing content, leading to degraded perceptual quality. On the other hand, recent generative TSE models typically synthesize high-quality speech with diffusion, but require numerous iterative steps resulting in high computational costs and latency. We propose Mask2Flow-TSE, a two-stage framework combining the strengths of both paradigms. We introduce the deletion/insertion (D/I) proportion, an analytical tool that reveals early flow steps predominantly remove signal components rather than synthesize them. Based on this finding, we decouple deletion from insertion: a masking-based module handles the deletion-dominant early steps, while a single flow-matching step performs the remaining insertion for high-quality reconstruction. Specifically, the first stage uses lightweight convolution for the masking module, while the second stage employs a Diffusion Transformer (DiT) adapted for TSE with speaker conditioning. Unlike prior approaches that start from Gaussian noise, our method starts from the masked spectrogram, enabling high-quality reconstruction in a single inference step. Experiments show that Mask2Flow-TSE produces high-quality extractions with only 85M parameters and one-step inference, while preserving clean single-speaker inputs with minimal degradation.


翻译:暂无翻译

0
下载
关闭预览

相关内容

ICML 2026教程:用KET统一Attention与Diffusion
专知会员服务
12+阅读 · 7月10日
【CVPR2022】语言作为查询的参考视频目标分割框架
专知会员服务
10+阅读 · 2022年4月27日
多语言语音识别声学模型建模方法最新进展
专知会员服务
36+阅读 · 2022年2月7日
FlowQA: Grasping Flow in History for Conversational Machine Comprehension
专知会员服务
35+阅读 · 2019年10月18日
视频目标检测:Flow-based
极市平台
22+阅读 · 2019年5月27日
近期语音类前沿论文
深度学习每日摘要
14+阅读 · 2019年3月17日
基于Tacotron模型的语音合成实践
深度学习每日摘要
15+阅读 · 2018年12月25日
原创 | Attention Modeling for Targeted Sentiment
黑龙江大学自然语言处理实验室
25+阅读 · 2017年11月5日
国家自然科学基金
122+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
《最强大的军事网状网络》
专知会员服务
0+阅读 · 今天14:29
《预测陆军征兵任务分配》110页
专知会员服务
1+阅读 · 今天14:21
分层反无人机系统发展新趋势
专知会员服务
9+阅读 · 9月3日
何为协作武器?
专知会员服务
10+阅读 · 9月1日
相关VIP内容
ICML 2026教程:用KET统一Attention与Diffusion
专知会员服务
12+阅读 · 7月10日
【CVPR2022】语言作为查询的参考视频目标分割框架
专知会员服务
10+阅读 · 2022年4月27日
多语言语音识别声学模型建模方法最新进展
专知会员服务
36+阅读 · 2022年2月7日
FlowQA: Grasping Flow in History for Conversational Machine Comprehension
专知会员服务
35+阅读 · 2019年10月18日
相关基金
国家自然科学基金
122+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员