Dialogue Enhancement (DE) enables the rebalancing of dialogue and background sounds to fit personal preferences and needs in the context of broadcast audio. When individual audio stems are unavailable from production, Dialogue Separation (DS) can be applied to the final audio mixture to obtain estimates of these stems. This work focuses on Preferred Loudness Differences (PLDs) between dialogue and background sounds. While previous studies determined the PLD through a listening test employing original stems from production, stems estimated by DS are used in the present study. In addition, a larger variety of signal classes is considered. PLDs vary substantially across individuals (average interquartile range: 5.7 LU). Despite this variability, PLDs are found to be highly dependent on the signal type under consideration, and it is shown that median PLDs can be predicted using objective intelligibility metrics. Two existing baseline prediction methods - intended for use with original stems - displayed a Mean Absolute Error (MAE) of 7.5 LU and 5 LU, respectively. A modified baseline (MAE: 3.2 LU) and an alternative approach (MAE: 2.5 LU) are proposed. Results support the viability of processing final broadcast mixtures with DS and offering an alternative remixing that accounts for median PLDs.
翻译:对话增强(DE)技术可在广播音频场景中重新平衡对话与背景声音,以适应用户的个性化偏好与需求。当无法获取制作环节的独立音频音轨时,可对最终音频混合信号应用对话分离(DS)技术来估算这些音轨。本研究聚焦于对话与背景声音之间的偏好响度差异(PLD)。此前研究通过使用制作环节原始音轨开展听力测试来确定PLD,而本研究则采用DS估算的音轨。此外,本研究还考虑了更广泛的信号类别。PLD在不同个体间存在显著差异(平均四分位距:5.7 LU)。尽管存在这种变异性,但研究发现PLD高度依赖于所评估的信号类型,并证明可通过客观清晰度指标预测中位数PLD。两种旨在与原始音轨配合使用的现有基线预测方法分别呈现7.5 LU和5 LU的平均绝对误差(MAE)。本研究提出了一种改进的基线方法(MAE:3.2 LU)和一种替代方法(MAE:2.5 LU)。研究结果支持使用DS处理最终广播混合信号的可行性,并验证了可考虑中位数PLD的替代混合方案的实用价值。