Introducing Shieldstral.
Summary
Shieldstral introduces a 3B open-weights multimodal safety classifier designed for content moderation. It uses a policy-adaptive, question-answering approach to produce calibrated safety scores for text and images without retraining, and is released under Apache 2.0 with weights hosted publicly. The article covers its architecture, benchmarks against larger models, and implementation details within Mistral's Forge platform.