Facial landmark detection is an essential technology for driver status tracking and has been in demand for real-time estimations. As a landmark coordinate prediction, heatmap-based methods are known to achieve a high accuracy, and Lite-HRNet can achieve a fast estimation. However, with Lite-HRNet, the problem of a heavy computational cost of the fusion block, which connects feature maps with different resolutions, has yet to be solved. In addition, the strong output module used in HRNetV2 is not applied to Lite-HRNet. Given these problems, we propose a novel architecture called Lite-HRNet Plus. Lite-HRNet Plus achieves two improvements: a novel fusion block based on a channel attention and a novel output module with less computational intensity using multi-resolution feature maps. Through experiments conducted on two facial landmark datasets, we confirmed that Lite-HRNet Plus further improved the accuracy in comparison with conventional methods, and achieved a state-of-the-art accuracy with a computational complexity with the range of 10M FLOPs.
翻译:人脸关键点检测是驾驶员状态追踪中的关键技术,对实时估计有较高需求。作为关键点坐标预测方法,基于热图的方法虽能实现高精度,而Lite-HRNet则可提供快速估计。然而,Lite-HRNet中连接不同分辨率特征图的融合块存在计算成本过高的问题,且HRNetV2中使用的强输出模块并未应用于Lite-HRNet。针对这些问题,我们提出了一种名为Lite-HRNet Plus的新型架构。该架构实现了两项改进:基于通道注意力的新型融合块,以及利用多分辨率特征图实现较低计算强度的新型输出模块。通过在两个人脸关键点数据集上的实验验证,我们确认Lite-HRNet Plus相较于传统方法进一步提升了精度,并在计算复杂度为10M FLOPs范围内达到了最优精度。