Decoding infant cry causes remains challenging for healthcare monitoring due to short nonstationary signals, limited annotations, and strong domain shifts across infants and datasets. We propose a compact acoustic framework that fuses mel-frequency cepstral coefficients (MFCCs), short-time Fourier transform (STFT) features, and fundamental-frequency (F0) contours within a multi-branch convolutional neural network (CNN) encoder, and models temporal dynamics using an enhanced Legendre Memory Unit (LMU). Compared to LSTMs, the LMU backbone provides stable sequence modeling with substantially fewer recurrent parameters, supporting efficient deployment. To improve cross-dataset generalization, we introduce calibrated posterior ensemble fusion with entropy-gated weighting to preserve domain-specific expertise while mitigating dataset bias. Experiments on Baby2020 and Baby Crying demonstrate improved macro-F1 under cross-domain evaluation, along with leakage aware splits and real-time feasibility for on-device monitoring.
翻译:解码婴儿啼哭原因对医疗监测仍具挑战,主要源于短时非平稳信号、标注稀疏以及跨婴儿与数据集的强域偏移。我们提出一种紧凑声学框架,该框架在多分支卷积神经网络编码器中融合梅尔频率倒谱系数、短时傅里叶变换特征与基频轮廓,并采用增强型勒让德记忆单元建模时间动态。相较于LSTM,LMU主干以显著更少的循环参数实现稳定序列建模,支撑高效部署。为提升跨数据集泛化能力,我们引入带熵门控权重的校准后验集成融合,在保留领域特异性知识的同时减轻数据集偏差。在Baby2020与Baby Crying数据集上的跨域评估显示宏F1得分提升,同时实现了泄漏感知分割与面向设备监测的实时可行性。