FineGen: A VLM-based Multi-Agent Framework for Fine-Grained Image-Text Dataset Construction
This work provides a scalable solution for generating fine-grained hard negatives, which is crucial for improving fine-grained perception in vision-language models.
FineGen addresses the scarcity of hard negative samples in vision-language datasets by introducing a VLM-based multi-agent framework that automatically constructs high-quality hard negatives. The resulting FineGen-100K dataset achieves a 96.7% attribute validity rate and yields a +14.4% accuracy improvement on hard samples in the FG-OVD benchmark.
The scarcity of hard negative samples in current vision-language datasets significantly hinders fine-grained perception. To address this, we propose FineGen, a VLM-based Multi-Agent framework for automated dataset construction. By employing a collaborative Generation-Verification-Correction pipeline with a closed-loop feedback mechanism, FineGen ensures synthesized hard negatives are semantically valid yet strictly contradictory to visual content. Applying this to ImageNet, we construct FineGen-100K, a hierarchical dataset containing over 147,000 attribute-specific hard negatives with a rigorous 1:10 positive-to-negative ratio. Extensive evaluations confirm a 96.7% attribute validity rate. Crucially, downstream validation on the FG-OVD benchmark shows that fine-tuning on FineGen-100K yields a substantial +14.4% accuracy improvement on hard samples, significantly outperforming state-of-the-art methods.