To obtain high-quality positron emission tomography (PET) images while minimizing radiation exposure, various methods have been proposed for reconstructing standard-dose PET (SPET) images from low-dose PET (LPET) sinograms directly. However, current methods often neglect boundaries during sinogram-to-image reconstruction, resulting in high-frequency distortion in the frequency domain and diminished or fuzzy edges in the reconstructed images. Furthermore, the convolutional architectures, which are commonly used, lack the ability to model long-range non-local interactions, potentially leading to inaccurate representations of global structures. To alleviate these problems, we propose a transformer-based model that unites triple domains of sinogram, image, and frequency for direct PET reconstruction, namely TriDo-Former. Specifically, the TriDo-Former consists of two cascaded networks, i.e., a sinogram enhancement transformer (SE-Former) for denoising the input LPET sinograms and a spatial-spectral reconstruction transformer (SSR-Former) for reconstructing SPET images from the denoised sinograms. Different from the vanilla transformer that splits an image into 2D patches, based specifically on the PET imaging mechanism, our SE-Former divides the sinogram into 1D projection view angles to maintain its inner-structure while denoising, preventing the noise in the sinogram from prorogating into the image domain. Moreover, to mitigate high-frequency distortion and improve reconstruction details, we integrate global frequency parsers (GFPs) into SSR-Former. The GFP serves as a learnable frequency filter that globally adjusts the frequency components in the frequency domain, enforcing the network to restore high-frequency details resembling real SPET images. Validations on a clinical dataset demonstrate that our TriDo-Former outperforms the state-of-the-art methods qualitatively and quantitatively.
翻译:为在最大化减少辐射暴露的同时获得高质量正电子发射断层扫描(PET)图像,研究者提出了多种方法,通过低剂量PET(LPET)正弦图直接重建标准剂量PET(SPET)图像。然而,现有方法在正弦图到图像的重建过程中常忽略边界信息,导致频域中出现高频畸变,重建图像边缘模糊或缺失。此外,常用的卷积架构缺乏建模长程非局部交互的能力,可能导致全局结构表征不准确。为解决这些问题,我们提出一种基于Transformer的模型——TriDo-Former,该模型融合正弦图、图像和频率三个域以实现直接PET重建。具体而言,TriDo-Former由两个级联网络组成:正弦图增强Transformer(SE-Former)用于对输入LPET正弦图进行去噪,以及空间-频谱重建Transformer(SSR-Former)用于从去噪后的正弦图重建SPET图像。与将图像分割为二维补丁的常规Transformer不同,我们的SE-Former基于PET成像机制,将正弦图分割为一维投影视角,以在去噪过程中保持其内部结构,从而防止正弦图中的噪声传播至图像域。此外,为缓解高频畸变并改善重建细节,我们在SSR-Former中集成了全局频率解析器(GFPs)。GFP作为可学习频率滤波器,可在频域中全局调整频率分量,迫使网络恢复接近真实SPET图像的高频细节。在临床数据集上的验证表明,我们的TriDo-Former在定性和定量性能上均优于现有最优方法。