Ghost imaging reconstructs spatial information from a single-pixel bucket detector by correlating structured illumination patterns with scalar intensity measurements. While deep learning approaches have achieved promising results on static scenes, two critical limitations remain unaddressed: existing architectures fail to exploit temporal coherence across frames, leaving dynamic ghost imaging largely unsolved, and they assume additive Gaussian noise models that do not reflect the true Poissonian statistics of real single-photon hardware. We present DynGhost (Dynamic Ghost Imaging Transformer), a transformer architecture that addresses both limitations through alternating spatial and temporal attention blocks. Our quantum-aware training framework, based on physically accurate detector simulations (SNSPDs, SPADs, SiPMs) and Anscombe variance-stabilizing normalization, resolves the distribution shift that causes classical models to fail under realistic hardware constraints. Experiments across multiple benchmarks demonstrate that DynGhost outperforms both traditional reconstruction methods and existing deep learning architectures, with particular gains in dynamic and photon-starved settings.
翻译:鬼成像通过将结构化照明模式与标量强度测量值相关联,从单像素桶探测器重建空间信息。尽管深度学习方法已在静态场景中取得令人瞩目的成果,但仍存在两个关键局限:现有架构无法利用帧间的时间相干性,导致动态鬼成像问题尚未解决;同时,它们假设了加性高斯噪声模型,这未能反映真实单光子硬件的泊松统计特性。我们提出DynGhost(动态鬼成像Transformer),这是一种通过交替执行空间注意力块与时间注意力块来同时解决上述两种局限的Transformer架构。基于物理精确的探测器仿真(SNSPD、SPAD、SiPM)以及Anscombe方差稳定归一化,我们构建了量子感知训练框架,解决了经典模型在现实硬件约束下失效的分布偏移问题。跨多个基准的实验表明,DynGhost在动态与光子匮乏场景下均优于传统重建方法及现有深度学习架构,尤其展现出显著优势。