Transformer-based models have significantly advanced natural language processing and computer vision in recent years. However, due to the irregular and disordered structure of point cloud data, transformer-based models for 3D deep learning are still in their infancy compared to other methods. In this paper we present Point Cross-Attention Transformer (PointCAT), a novel end-to-end network architecture using cross-attentions mechanism for point cloud representing. Our approach combines multi-scale features via two seprate cross-attention transformer branches. To reduce the computational increase brought by multi-branch structure, we further introduce an efficient model for shape classification, which only process single class token of one branch as a query to calculate attention map with the other. Extensive experiments demonstrate that our method outperforms or achieves comparable performance to several approaches in shape classification, part segmentation and semantic segmentation tasks.
翻译:近年来,基于Transformer的模型显著推进了自然语言处理与计算机视觉领域的发展。然而,由于点云数据具有不规则和无序的结构特点,相比其他方法,基于Transformer的3D深度学习模型仍处于发展初期。本文提出点云交叉注意力Transformer(PointCAT),这是一种利用交叉注意力机制进行点云表示的新型端到端网络架构。该方法通过两个独立的交叉注意力Transformer分支融合多尺度特征。为降低多分支结构带来的计算量增长,我们进一步引入了一种高效的形状分类模型,该模型仅将一个分支中的单一类令牌作为查询,与另一分支计算注意力图。大量实验表明,在形状分类、部件分割和语义分割任务中,我们的方法优于或达到与多种方法相当的性能。