Optical microrobots actuated by optical tweezers (OT) are important for cell manipulation and microscale assembly, but their autonomous operation depends on accurate 3D perception. Developing such perception systems is challenging because large-scale, high-quality microscopy datasets are scarce, owing to complex fabrication processes and labor-intensive annotation. Although generative AI offers a promising route for data augmentation, existing generative adversarial network (GAN)-based methods struggle to reproduce key optical characteristics, particularly depth-dependent diffraction and defocus effects. To address this limitation, we propose Du-FreqNet, a dual-control, frequency-aware diffusion model for physically consistent microscopy image synthesis. The framework features two independent ControlNet branches to encode microrobot 3D point clouds and depth-specific mesh layers, respectively. We introduce an adaptive frequency-domain loss that dynamically reweights high- and low-frequency components based on the distance to the focal plane. By leveraging differentiable FFT-based supervision, Du-FreqNet captures physically meaningful frequency distributions often missed by pixel-space methods. Trained on a limited dataset (e.g., 80 images per pose), our model achieves controllable, depth-dependent image synthesis, improving SSIM by 20.7% over baselines. Extensive experiments demonstrate that Du-FreqNet generalizes effectively to unseen poses and significantly enhances downstream tasks, including 3D pose and depth estimation, thereby facilitating robust closed-loop control in microrobotic systems.
翻译:光镊驱动的光驱动微机器人在细胞操作和微尺度组装中至关重要,但其自主运行依赖于精确的三维感知。由于复杂的制造工艺和人力密集的标注工作,大规模、高质量显微数据集的稀缺使得开发此类感知系统充满挑战。尽管生成式人工智能为数据增强提供了有前景的途径,但现有的基于生成对抗网络的方法难以复现关键光学特性,尤其是深度相关的衍射和离焦效应。针对这一局限,我们提出了Du-FreqNet——一种用于物理一致性显微图像合成的双控频率感知扩散模型。该框架包含两个独立的ControlNet分支,分别编码微机器人三维点云和深度特定网格层。我们引入了一种自适应频域损失函数,可根据到焦平面的距离动态地重新加权高频和低频分量。通过利用可微分的基于FFT的监督,Du-FreqNet捕获了像素空间方法常遗漏的具有物理意义的频率分布。在有限数据集(例如每姿态80张图像)上训练后,我们的模型实现了可控的、深度相关的图像合成,相比基线方法,结构相似性指数提高了20.7%。大量实验表明,Du-FreqNet能够有效泛化至未见姿态,并显著增强下游任务(包括三维姿态和深度估计),从而促进微机器人系统中稳健的闭环控制。