CLCVJul 28

Shieldstral

arXiv:2607.2585728.2
Predicted impact top 2% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For AI safety researchers and deployers, it provides a small, efficient model that achieves high performance on diverse safety tasks, enabling broader adoption.

Shieldstral is a 3B-parameter multimodal safety classifier that matches or outperforms models nearly 7x its size on text safety benchmarks and sets a new state of the art on multimodal safety classification.

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes