To address the increasing need for efficient and accurate content moderation, we propose an efficient and lightweight deep classification ensemble structure. Our approach is based on a combination of simple visual features, designed for high-accuracy classification of violent content with low false positives. Our ensemble architecture utilizes a set of lightweight models with narrowed-down color features, and we apply it to both images and videos. We evaluated our approach using a large dataset of explosion and blast contents and compared its performance to popular deep learning models such as ResNet-50. Our evaluation results demonstrate significant improvements in prediction accuracy, while benefiting from 7.64x faster inference and lower computation cost. While our approach is tailored to explosion detection, it can be applied to other similar content moderation and violence detection use cases as well. Based on our experiments, we propose a "think small, think many" philosophy in classification scenarios. We argue that transforming a single, large, monolithic deep model into a verification-based step model ensemble of multiple small, simple, and lightweight models with narrowed-down visual features can possibly lead to predictions with higher accuracy.
翻译:为应对日益增长的高效且精准的内容审核需求,我们提出了一种高效轻量的深度学习分类集成架构。该方法基于简单视觉特征的组合,旨在以低误报率实现暴力内容的高精度分类。所提出的集成架构采用一组窄化颜色特征的轻量模型,并同时应用于图像与视频。我们使用爆炸及爆破内容的大规模数据集对方法进行评估,并将其性能与ResNet-50等热门深度学习模型进行对比。评估结果表明,该方法在预测精度上取得显著提升,同时推理速度加快7.64倍且计算成本降低。尽管本方法针对爆炸检测场景设计,但同样可推广至其他类似的内容审核与暴力检测应用场景。基于实验结果,我们提出分类场景中的"以小见多"理念:将单一庞大且整体式的深度模型,转化为由多个窄化视觉特征的轻量简单模型构成的验证式逐步集成模型,可能获得更高精度的预测结果。