CVJun 16

Million-scale multimodal pollen microscopy with expert-guided foundation models

arXiv:2606.178097.1
Predicted impact top 69% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For palynologists and ecologists, this provides a large-scale benchmark for automated pollen identification that generalizes across scanners and regions, addressing a key bottleneck in aerobiology and biodiversity monitoring.

The authors created a million-scale multimodal pollen microscopy dataset (Pollen AI Atlas) with 1.5M grain detections and machine-generated morphological captions, achieving 99.6% proposal precision and 88.16% top-1 accuracy in baseline benchmarks. Cross-regional retrieval showed caption embeddings robust when image similarity degraded (mAP@20 0.811 vs 0.262).

Automated pollen identification from microscopy remains a bottleneck in aerobiology, palaeoecology and biodiversity monitoring, because scalable systems must generalise across specimen preparation, scanner settings and geographic origins while retaining palynological interpretability. To address this gap, we present a million-scale multimodal pollen microscopy resource, Pollen AI Atlas, assembled from pure-species whole-slide bright-field images spanning four geographic origins, four scanner settings and 46 taxon labels across 31 botanical families. Seeded by one manually selected exemplar per source slide, token-level mining and filtering produced 1,511,390 released grain detections with 99.6\% proposal precision in expert-curated test regions. Each detection was paired with machine-generated grain-level morphological captions from five open-weight vision-language models, guided by expert-verified palynological anchors, yielding structured descriptions of aperture systems, wall ornamentation, shape and size. Among the evaluated models, Gemma4 provided the most controlled primary caption set, combining tight length control, no leakage and the strongest text-retrieval performance. Baseline benchmarks with frozen visual features reached 88.16\% top-1 accuracy, while cross-regional retrieval showed that caption-derived text embeddings remained robust when image similarity degraded (mAP@20 0.811 versus 0.262). Released data, annotations, captions, splits, code, and weights provide a benchmark for pollen recognition, cross-regional domain adaptation and domain-specific multimodal microscopy learning.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes