Multi-pitch estimation (MPE) typically predicts which pitches are active in a mixture, but not which instrument or source produced them. This paper investigates a lightweight slot-attention framework for multi-instrument MPE (MI-MPE), where a mixture CQT is mapped to an unordered set of source-like pitch maps. The model uses permutation-invariant Hungarian matching to avoid fixed output semantics and treats the number of slots as an upper bound on the number of active sources. We further study two modular extensions: a self-supervised timbre encoder that provides training-time targets for slot-level timbre embeddings, and a polyphony branch that regularizes the pitch density of mixture- and slot-level predictions. Experiments show that Hungarian matching substantially improves instrument family decomposition on URMP. Stem-level prediction remains more challenging: timbre and polyphony supervision improve selected configurations, but do not consistently resolve source assignment. The results suggest that slot-based architectures are a promising direction for source-aware MPE, while highlighting the need to couple auxiliary musical cues to slot identity more carefully.
翻译:多音高估计(MPE)通常预测混合音频中哪些音高是活跃的,但不会预测产生它们的乐器或声源。本文研究了一种用于多乐器MPE(MI-MPE)的轻量级槽注意力框架,该框架将混合音频的CQT映射为无序的源类音高图集合。模型采用置换不变的匈牙利匹配以避免固定的输出语义,并将槽的数量视为活跃声源数量的上界。我们进一步研究了两种模块化扩展:一种自监督音色编码器,为槽级音色嵌入提供训练时目标;以及一个复调分支,用于规范化混合级和槽级预测的音高密度。实验表明,匈牙利匹配显著提升了URMP数据集上的乐器族分解效果。音轨级预测仍然更具挑战性:音色和复调监督改进了特定配置,但未能一致地解决声源分配问题。结果表明,基于槽的架构是声源感知MPE的一个有前景的方向,同时也凸显了将辅助音乐线索与槽标识更仔细地耦合的必要性。