Machine learning and deep learning have been used extensively to classify physical surfaces through images and time-series contact data. However, these methods rely on human expertise and entail the time-consuming processes of data and parameter tuning. To overcome these challenges, we propose an easily implemented framework that can directly handle heterogeneous data sources for classification tasks. Our data-versus-data approach automatically quantifies distinctive differences in distributions in a high-dimensional space via kernel two-sample testing between two sets extracted from multimodal data (e.g., images, sounds, haptic signals). We demonstrate the effectiveness of our technique by benchmarking against expertly engineered classifiers for visual-audio-haptic surface recognition due to the industrial relevance, difficulty, and competitive baselines of this application; ablation studies confirm the utility of key components of our pipeline. As shown in our open-source code, we achieve 97.2% accuracy on a standard multi-user dataset with 108 surface classes, outperforming the state-of-the-art machine-learning algorithm by 6% on a more difficult version of the task. The fact that our classifier obtains this performance with minimal data processing in the standard algorithm setting reinforces the powerful nature of kernel methods for learning to recognize complex patterns.
翻译:机器学习和深度学习已广泛应用于通过图像和时间序列接触数据对物理表面进行分类。然而,这些方法依赖人类专业知识,且数据与参数调优过程耗时。为克服上述挑战,我们提出了一种易于实现的框架,可直接处理异构数据源以完成分类任务。我们的数据对数据方法通过核双样本检验,自动量化从多模态数据(如图像、声音、触觉信号)中提取的两组数据在高维空间中的分布差异。鉴于该应用在工业中的相关性、难度及竞争性基线,我们通过与专家设计的分类器在视听触表面识别任务上进行基准测试,验证了所提技术的有效性;消融研究进一步证实了我们流程中关键组件的实用性。如开源代码所示,我们在包含108个表面类别的标准多用户数据集上实现了97.2%的准确率,在任务更困难版本上比当前最先进的机器学习算法高出6%。我们的分类器在标准算法设置下以最小数据处理量获得此性能,这强化了核方法在识别复杂模式学习中的强大特性。