Feature visualization is used to visualize learned features for black box machine learning models. Our approach explores an altered training process to improve interpretability of the visualizations. We argue that by using background removal techniques as a form of robust training, a network is forced to learn more human recognizable features, namely, by focusing on the main object of interest without any distractions from the background. Four different training methods were used to verify this hypothesis. The first used unmodified pictures. The second used a black background. The third utilized Gaussian noise as the background. The fourth approach employed a mix of background removed images and unmodified images. The feature visualization results show that the background removed images reveal a significant improvement over the baseline model. These new results displayed easily recognizable features from their respective classes, unlike the model trained on unmodified data.
翻译:特征可视化用于可视化黑箱机器学习模型中学习到的特征。我们的方法探索了一种修改后的训练过程,以提高可视化结果的可解释性。我们认为,通过将背景移除技术作为一种鲁棒训练形式,网络被迫学习更易被人类识别的特征,即通过聚焦于主要感兴趣对象,而不受背景的任何干扰。我们使用了四种不同的训练方法来验证这一假设。第一种使用未修改的图片。第二种使用黑色背景。第三种将高斯噪声作为背景。第四种方法结合了背景移除图像和未修改图像。特征可视化结果显示,背景移除图像相对于基线模型有显著提升。这些新结果展示了各自类别中易于识别的特征,与基于未修改数据训练的模型不同。