Robotics & Physical AIAug 5, 2026
Mistral releases Shieldstral, a 3B open-weights policy-adaptive safety classifier
On August 4, 2026, Mistral released Shieldstral 1.0, a 3-billion-parameter open-weights moderation model (Apache 2.0, on Hugging Face) that judges text and images against policies written in plain language at inference time instead of fixed harm categories. Mistral says it matches guard models up to seven times its size on text safety and runs on a single 16GB GPU across 12 languages.
What it means Guardrails you can rewrite in plain language without retraining — and self-host on commodity hardware — change the cost calculus for anyone shipping user-facing AI.
Where it came from Mistral AI