Automated pollen identification from microscopy remains a bottleneck in aerobiology, palaeoecology and biodiversity monitoring, because scalable systems must generalise across specimen preparation, scanner settings and geographic origins while retaining palynological interpretability. To address this gap, we present a million-scale multimodal pollen microscopy resource, Pollen AI Atlas, assembled from pure-species whole-slide bright-field images spanning four geographic origins, four scanner settings and 46 taxon labels across 31 botanical families. Seeded by one manually selected exemplar per source slide, token-level mining and filtering produced 1,511,390 released grain detections with 99.6\% proposal precision in expert-curated test regions. Each detection was paired with machine-generated grain-level morphological captions from five open-weight vision-language models, guided by expert-verified palynological anchors, yielding structured descriptions of aperture systems, wall ornamentation, shape and size. Among the evaluated models, Gemma4 provided the most controlled primary caption set, combining tight length control, no leakage and the strongest text-retrieval performance. Baseline benchmarks with frozen visual features reached 88.16\% top-1 accuracy, while cross-regional retrieval showed that caption-derived text embeddings remained robust when image similarity degraded (mAP@20 0.811 versus 0.262). Released data, annotations, captions, splits, code, and weights provide a benchmark for pollen recognition, cross-regional domain adaptation and domain-specific multimodal microscopy learning.
翻译:自动花粉显微鉴定仍是空气生物学、古生态学和生物多样性监测领域的瓶颈,因为可扩展系统需在样本制备方式、扫描仪设置和地理来源上实现泛化,同时保持孢粉学可解释性。为此,我们构建了百万级多模态花粉显微资源——花粉AI图谱(Pollen AI Atlas),该资源包含来自31个植物科系的纯种全切片明场图像,涵盖四个地理来源、四种扫描仪设置和46个分类标签。通过每张源切片人工选取一个示范样本作为种子,经令牌级挖掘与过滤,在专家标注测试区域以99.6%的提案精度获得1,511,390个释放的花粉检测结果。每个检测结果均配有机生成的颗粒级形态学描述文本,该描述基于五种开源视觉语言模型,并辅以专家验证的孢粉学锚点引导,最终形成包含萌发孔系统、纹饰、形状和尺寸的结构化描述。在评估模型中,Gemma4提供了最受控的主描述文本集,兼具严格长度控制、无信息泄露和最强文本检索性能。基于冻结视觉特征的基线基准测试达到88.16%的Top-1准确率,而跨区域检索实验表明,当图像相似度下降时,基于描述文本的嵌入仍保持稳健(mAP@20从0.262提升至0.811)。所发布的数据、标注、描述文本、数据集划分方案、代码及模型权重,为花粉识别、跨区域领域适应及领域特定多模态显微学习提供了基准。