The proliferation of machine learning (ML) has drawn unprecedented interest in the study of various multimedia contents such as text, image, audio and video, among others. Consequently, understanding and learning ML-based representations have taken center stage in knowledge discovery in intelligent multimedia research and applications. Nevertheless, the black-box nature of contemporary ML, especially in deep neural networks (DNNs), has posed a primary challenge for ML-based representation learning. To address this black-box problem, the studies on interpretability of ML have attracted tremendous interests in recent years. This paper presents a survey on recent advances and future prospects on interpretability of ML, with several application examples pertinent to multimedia computing, including text-image cross-modal representation learning, face recognition, and the recognition of objects. It is evidently shown that the study of interpretability of ML promises an important research direction, one which is worth further investment in.
翻译:机器学习的普及引起了人们对文本、图像、音频与视频等多种多媒体内容研究的空前关注。因此,在智能多媒体研究与应用中,理解并学习基于机器学习的表征已成为知识发现的核心议题。然而,当代机器学习(尤其是深度神经网络)的"黑箱"特性对基于机器学习的表征学习构成了主要挑战。为解决这一黑箱问题,近年来机器学习可解释性研究已引发极大兴趣。本文综述了机器学习可解释性的最新进展与未来前景,并给出了多个与多媒体计算相关的应用实例,包括文本-图像跨模态表征学习、人脸识别与物体识别等。研究结果表明,机器学习可解释性研究预示着重要的研究方向,值得进一步投入。