Abstract reasoning problems pose challenges to the perception and cognition abilities of AI algorithms, demanding deeper pattern recognition and inductive reasoning beyond mere identification of explicit image features. In this study, we introduce PMoC, a probabilistic model tailored for the Bongard-Logo problem, achieving high reasoning accuracy through the construction of an conditional probabilistic model. Additionally, we have designed the Pose-Transformer, an enhanced Transformer-Encoder specifically crafted for complex abstract reasoning tasks, including Bongard-Logo, RAVEN, I-RAVEN, and PGM. Inspired by the pose matrix in capsule networks, Pose-Transformer strengthens the focus on positional relationships between local features when processing image data. When combined with PMoC, it can further enhance reasoning accuracy. Our Pose-Transformer effectively addresses reasoning difficulties associated with changes in the position of abstract entities, outperforming previous models on RAVEN's OIG, D3$\times$3 subsets, and the PGM dataset. Finally, considering the deployment difficulties arising from the large number of Pose-Transformer parameters, this paper presents a lightweight version, Straw-Pose-Transformer, which maintains performance while significantly reducing the parameter count. This study contributes to enhancing AI capabilities in abstract reasoning and cognitive pattern recognition.
翻译:抽象推理问题对人工智能算法的感知与认知能力提出了挑战,要求其不仅识别显式图像特征,还需进行更深层次的模式识别与归纳推理。本研究针对Bongard-Logo问题提出了PMoC概率模型,通过构建条件概率模型实现了较高的推理准确率。此外,我们设计了Pose-Transformer——一种专为复杂抽象推理任务(包括Bongard-Logo、RAVEN、I-RAVEN和PGM)优化的增强型Transformer-Encoder。受胶囊网络中姿态矩阵的启发,Pose-Transformer在处理图像数据时加强了对局部特征间位置关系的关注。当与PMoC结合时,能进一步提升推理精度。我们的Pose-Transformer有效解决了抽象实体位置变化带来的推理困难,在RAVEN的OIG、D3$\times$3子集及PGM数据集上超越了先前模型。最后,针对Pose-Transformer参数量过大导致的部署难题,本文提出了轻量化版本Straw-Pose-Transformer,在保持性能的同时显著减少了参数量。本研究为提升人工智能在抽象推理与认知模式识别方面的能力作出了贡献。