This study presents a novel approach to bone age assessment (BAA) using a multi-view, multi-task classification model based on the Sauvegrain method. A straightforward solution to automating the Sauvegrain method, which assesses a maturity score for each landmark in the elbow and predicts the bone age, is to train classifiers independently to score each region of interest (RoI), but this approach limits the accessible information to local morphologies and increases computational costs. As a result, this work proposes a self-accumulative vision transformer (SAT) that mitigates anisotropic behavior, which usually occurs in multi-view, multi-task problems and limits the effectiveness of a vision transformer, by applying token replay and regional attention bias. A number of experiments show that SAT successfully exploits the relationships between landmarks and learns global morphological features, resulting in a mean absolute error of BAA that is 0.11 lower than that of the previous work. Additionally, the proposed SAT has four times reduced parameters than an ensemble of individual classifiers of the previous work. Lastly, this work also provides informative implications for clinical practice, improving the accuracy and efficiency of BAA in diagnosing abnormal growth in adolescents.
翻译:本研究提出了一种基于Sauvegrain方法的多视图、多任务分类模型,用于骨骼年龄评估(BAA)的创新方法。实现Sauvegrain方法自动化的直接解决方案是独立训练分类器以对每个感兴趣区域(RoI)进行评分,但这种方法将可用信息限制于局部形态,并增加了计算成本。为此,本文提出了一种自累加视觉Transformer(SAT),通过应用令牌重放和区域注意力偏置,有效缓解了多视图、多任务问题中常见的各向异性行为——该类行为会削弱视觉Transformer的效能。大量实验表明,SAT成功利用了各标志点间的关联性并学习全局形态特征,使得BAA的平均绝对误差比先前工作低0.11。此外,所提出的SAT参数量仅为先前工作中独立分类器集成模型的四分之一。最后,本研究还为临床实践提供了重要启示,通过提升诊断青少年生长异常时BAA的准确性与效率。