We introduce CatNet, an algorithm that effectively controls False Discovery Rate (FDR) and selects significant features in LSTM. CatNet employs the derivative of SHAP values to quantify the feature importance, and constructs a vector-formed mirror statistic for FDR control with the Gaussian Mirror algorithm. To avoid instability due to nonlinear or temporal correlations among features, we also propose a new kernel-based independence measure. CatNet performs robustly on different model settings with both simulated and real-world data, which reduces overfitting and improves interpretability of the model. Our framework that introduces SHAP for feature importance in FDR control algorithms and improves Gaussian Mirror can be naturally extended to other time-series or sequential deep learning models.
翻译:我们提出CatNet算法,该算法能够有效控制LSTM中的错误发现率(FDR)并筛选重要特征。CatNet利用SHAP值的导数量化特征重要性,结合高斯镜像算法构建向量形式的镜像统计量以实现FDR控制。为解决特征间非线性或时序相关性导致的算法不稳定性,我们进一步提出基于核函数的独立性度量新方法。在模拟数据和真实数据集上,CatNet在不同模型配置下均展现出稳健性能,能够有效降低过拟合风险并提升模型可解释性。本框架通过引入SHAP实现FDR控制算法中的特征重要性建模,并改进高斯镜像方法,可自然扩展至其他时间序列或序列深度学习模型。