To interpret deep neural networks, one main approach is to dissect the visual input and find the prototypical parts responsible for the classification. However, existing methods often ignore the hierarchical relationship between these prototypes, and thus can not explain semantic concepts at both higher level (e.g., water sports) and lower level (e.g., swimming). In this paper inspired by human cognition system, we leverage hierarchal information to deal with uncertainty: When we observe water and human activity, but no definitive action it can be recognized as the water sports parent class. Only after observing a person swimming can we definitively refine it to the swimming action. To this end, we propose HIerarchical Prototype Explainer (HIPE) to build hierarchical relations between prototypes and classes. HIPE enables a reasoning process for video action classification by dissecting the input video frames on multiple levels of the class hierarchy, our method is also applicable to other video tasks. The faithfulness of our method is verified by reducing accuracy-explainability trade off on ActivityNet and UCF-101 while providing multi-level explanations.
翻译:为了解释深度神经网络,一种主要方法是通过解剖视觉输入并找出负责分类的原型特征部分。然而,现有方法往往忽略这些原型之间的层次关系,因此无法同时解释高层级(如水上运动)和低层级(如游泳)的语义概念。受人类认知系统启发,本文利用层次化信息处理不确定性:当观察到水和人类活动但未确定具体动作时,可将其识别为水上运动父类别;只有观察到人游泳后,才能精确定位为游泳动作。为此,我们提出层次化原型解释器(HIPE),以构建原型与类别间的层次关系。HIPE通过从类别层次结构的多个层级解剖输入视频帧,实现视频动作分类的推理过程,该方法同样适用于其他视频任务。通过在ActivityNet和UCF-101数据集上降低准确率与可解释性之间的权衡,同时提供多层级解释,验证了本方法的忠实性。