Real-time intelligent detection and prediction of subjects' behavior particularly their movements or actions is critical in the ward. This approach offers the advantage of reducing in-hospital care costs and improving the efficiency of healthcare workers, which is especially true for scenarios at night or during peak admission periods. Therefore, in this work, we propose using computer vision (CV) and deep learning (DL) methods for detecting subjects and recognizing their actions. We utilize OpenPose as an accurate subject detector for recognizing the positions of human subjects in the video stream. Additionally, we employ AlphAction's Asynchronous Interaction Aggregation (AIA) network to predict the actions of detected subjects. This integrated model, referred to as PoseAction, is proposed. At the same time, the proposed model is further trained to predict 12 common actions in ward areas, such as staggering, chest pain, and falling down, using medical-related video clips from the NTU RGB+D and NTU RGB+D 120 datasets. The results demonstrate that PoseAction achieves the highest classification mAP of 98.72% ([email protected]). Additionally, this study develops an online real-time mode for action recognition, which strongly supports the clinical translation of PoseAction. Furthermore, using OpenPose's function for recognizing face key points, we also implement face blurring, which is a practical solution to address the privacy protection concerns of patients and healthcare workers. Nevertheless, the training data for PoseAction is currently limited, particularly in terms of label diversity. Consequently, the subsequent step involves utilizing a more diverse dataset (including general actions) to train the model's parameters for improved generalization.
翻译:摘要:在病房环境中,实时智能检测与预测对象行为(尤其是动作与移动)至关重要。该方法具有降低院内护理成本、提升医护人员工作效率的优势,在夜间或就诊高峰期场景中尤为显著。为此,本研究提出采用计算机视觉(CV)与深度学习(DL)方法实现患者检测及动作识别。我们采用OpenPose作为高精度人体检测器,用于识别视频流中人体目标的位置。同时,运用AlphAction的异步交互聚合(AIA)网络预测已检测目标的动作,由此提出集成模型PoseAction。该模型进一步利用NTU RGB+D与NTU RGB+D 120数据集中的医疗相关视频片段进行训练,可识别病房区域12种常见动作(如蹒跚、胸痛、跌倒等)。实验结果表明,PoseAction在[email protected]条件下分类mAP最高达98.72%。此外,本研究开发了在线实时动作识别模式,为PoseAction的临床转化提供有力支撑。通过利用OpenPose的人脸关键点识别功能,我们还实现了人脸模糊处理,为保护患者与医护人员隐私提供了实用解决方案。然而,当前PoseAction的训练数据仍存在标签多样性有限的局限。因此,后续工作将采用包含通用动作的更丰富数据集进行模型参数训练,以提升泛化能力。